Menu

Apache Spark course · Lesson 1 of 4

Apache Spark Architecture and Execution Model

How Apache Spark works end to end: driver and executors, cluster managers, the Catalyst optimiser, jobs, stages and tasks, shuffles, memory and adaptive execution.

  • Intermediate
  • Pillar guide
  • 2 min read
  • Updated Oct 2026
On this page
  1. The processes
  2. From code to a plan
  3. From plan to tasks
  4. Shuffles
  5. Joins
  6. Memory
  7. Adaptive Query Execution
  8. Where Spark runs
  9. Checkpoints

Understanding Spark’s architecture explains almost every performance behaviour you will see. This guide walks from the processes that make up an application to how a query becomes tasks.

The processes

  • Driver: runs your program, builds and optimises the plan, schedules tasks and collects results.
  • Executors: processes on worker machines that run tasks, cache data and store shuffle files. Each has a number of cores; each core runs one task at a time.
  • Cluster manager: allocates resources for executors (for example Kubernetes, YARN, standalone, or a managed platform’s own scheduler).

Parallelism is roughly the number of executors times cores per executor.

From code to a plan

DataFrame and SQL code produce a logical plan. Spark’s optimiser (Catalyst) rewrites it: pushing filters toward the data source, pruning unused columns, reordering joins and choosing physical operators. The result is a physical plan that you can inspect with explain().

Read: Transformations vs actions

From plan to tasks

An action submits work as one or more jobs. Each job is split into stages at shuffle boundaries, and each stage runs one task per partition.

Read: Jobs, stages and tasks · Practise: Explain jobs, stages and tasks

Shuffles

A shuffle redistributes data so rows with the same key meet on the same executor: map tasks write partitioned shuffle files, reduce tasks fetch them over the network. Shuffles are the most expensive part of most jobs and the place where skew appears.

Read: Partitions, shuffles and skew · Practise: Data skew, Partition count

Joins

Spark chooses a physical join strategy: broadcast hash join when one side is small, sort-merge join for two large sides, and others in specific cases. The choice often dominates job cost.

Read: Joins and join strategy

Memory

Executor memory is shared between execution (joins, aggregations, sorts) and storage (cached data). When execution memory runs out, Spark spills to disk, which is slower but keeps the job alive; very large partitions or skewed keys are the usual cause.

Adaptive Query Execution

At each shuffle boundary, AQE re-plans the rest of the query using real statistics: coalescing small partitions, switching to broadcast joins and splitting skewed join partitions. It is on by default in modern Spark.

Read: Adaptive Query Execution

Where Spark runs

Spark runs on many platforms, including managed services such as Databricks. The architecture is the same; the platform manages clusters, storage integration and scheduling.

Checkpoints

You understand Spark’s architecture when you can explain, for a given job, how many stages it has and why, which join strategy was used, where the shuffles are, and what the Spark UI would show if one key were skewed.

Revise with the Spark interview cheat sheet.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Defaults and plans in the linked guides were checked on Spark 4.2

Progress is saved in this browser only. No account needed.

Search
Filter by type