Menu

Airflow interview question · Question 1 of 1

How should Airflow retries and idempotency work together?

  • Medium
  • conceptual / scenario
  • ~7 min
  • High relevance
  • 2 min read
  • Updated Oct 2026

Short answer

Retries make Airflow rerun a failed task, possibly after it already wrote part of its output, so retries are only safe if the task is idempotent. Each task should process exactly its run's data interval, taken from the run context rather than the current time, and write by overwriting that interval's partition or upserting on a unique key inside a transaction. Then a retry, a manual clear or a backfill all produce the same final state.

On this page
  1. Detailed explanation
  2. Design rules
  3. Example settings
  4. Common mistakes

Detailed explanation

Rerun cause Why idempotency matters
Automatic retry The first attempt may have written partial output
Manual “clear” of a task Someone reruns after a fix
Backfill Old intervals are processed again

Design rules

  1. Deterministic input: use the run’s data interval (data_interval_start, data_interval_end).
  2. Replace, not append: partition overwrite, MERGE, or build-and-swap.
  3. Atomic writes: transactions or table formats with atomic commits.
  4. Guard side effects: record that a notification was sent for this run before sending again.
  5. Retry the right failures: transient errors with backoff; let schema errors fail.

Example settings

from datetime import timedelta

default_args = {
    "retries": 3,
    "retry_delay": timedelta(minutes=2),
    "retry_exponential_backoff": True,
    "max_retry_delay": timedelta(minutes=30),
}

Common mistakes

  1. Append-only loads with retries enabled.
  2. Using the current time instead of the data interval.
  3. Very high retry counts that hide a real failure for hours.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Concepts apply to Airflow 2.x and 3.x

Progress is saved in this browser only. No account needed.

Search
Filter by type