Menu

Apache Spark interview question · Question 2 of 5

Explain Spark jobs, stages and tasks.

  • Medium
  • conceptual
  • ~6 min
  • High relevance
  • 2 min read
  • Updated Oct 2026

Short answer

When you call an action, Spark creates a job to compute it. The job is divided into stages at shuffle boundaries, because each wide transformation needs all data for a key before continuing. Each stage runs as a set of tasks, one per partition, executed in parallel on executor cores. With Adaptive Query Execution enabled, one action may show up as several jobs, because Spark re-plans after each shuffle stage.

On this page
  1. Detailed explanation
  2. Example
  3. Driver versus executors
  4. Common mistakes

Detailed explanation

Concept Created by Count determined by
Job An action (count, write, show) One or more per action
Stage Shuffle boundaries in the plan Number of wide transformations + 1 (roughly)
Task A partition in a stage Number of partitions in that stage

Example

Reading a dataset with 8 input partitions, filtering, then groupBy("country").count() and show():

  1. Job: triggered by show().
  2. Stage 1: read + filter + partial aggregation, 8 tasks (one per input partition), ending with shuffle write.
  3. Stage 2: shuffle read + final aggregation, as many tasks as shuffle partitions (often fewer after AQE coalescing).

Driver versus executors

The driver plans and schedules; executors run tasks and store shuffle and cached data. A slow driver usually means too much collect() or too many tiny tasks to schedule.

Common mistakes

  1. Saying each transformation is a stage (only shuffles create new stages).
  2. Equating tasks with executors (tasks map to partitions; executors provide cores).

By Data Career Hub Editorial · Last reviewed Oct 2026 · Examples run on PySpark 4.2 in local mode; behaviour notes say where Spark 3.x differs

Progress is saved in this browser only. No account needed.

Search
Filter by type