Apache Spark interview questionsQuestion 3 of 5
Apache Spark interview question · Question 3 of 5
How does partition count affect Spark performance?
Short answer
Each partition becomes one task, so partition count sets parallelism. Too few partitions leave cores idle and make each task large enough to spill or run out of memory; too many tiny partitions waste time on scheduling and produce small output files. Aim for evenly sized partitions, often in the range of 100 to 200 MB, at least a few times the number of cores, and let Adaptive Query Execution coalesce shuffle partitions for you.
Detailed explanation
- Input partitions when reading files come mostly from file sizes and
spark.sql.files.maxPartitionBytes(128 MB by default). - Shuffle partitions come from
spark.sql.shuffle.partitions(200 by default), which AQE can coalesce. - Output files: each task writing to a location usually produces at least one file, so partition count also controls the small-files problem.
Rules of thumb
| Situation | Action |
|---|---|
| Tasks spill or run out of memory | More, smaller partitions |
| Thousands of tasks of a few MB | Fewer partitions; rely on AQE coalescing; compact inputs |
| Many tiny output files | coalesce before writing, or repartition by the output partition column |
| A few tasks much slower | Skew, not count: fix the key distribution |
repartition(n) does a full shuffle and can increase or decrease the count evenly; coalesce(n) only decreases it, merging partitions without a full shuffle, but may leave them uneven.
Common mistakes
- Treating skew as a partition-count problem.
repartition(1)before writing large data (one task does all the work).- Hard-coding 200 shuffle partitions for every job size.
Progress is saved in this browser only. No account needed.