Menu

PySpark interview question · Question 1 of 2

When should you avoid Python UDFs in PySpark?

  • Medium
  • conceptual / optimization
  • ~6 min
  • High relevance
  • 2 min read
  • Updated Oct 2026

Short answer

Avoid a Python UDF whenever built-in Spark functions can express the logic. A UDF moves data between the JVM and Python workers and is a black box to the optimiser, so it is slower and prevents optimisations such as predicate pushdown. If Python is genuinely needed, prefer a vectorised pandas UDF that processes batches through Apache Arrow, and keep any UDF pure, null-safe and tested.

On this page
  1. Detailed explanation
  2. Decision order
  3. Example
  4. Common mistakes

Detailed explanation

Built-in functions run in the JVM with code generation and are understood by the optimiser. A Python UDF adds a separate Python evaluation step in the physical plan. In Spark 4, regular Python UDFs can use Arrow for transfer by default when PyArrow is installed, which narrows the serialisation gap, but the function is still opaque to the optimiser and still executes Python.

Decision order

  1. Built-in function (when, regexp_extract, date_trunc, array functions).
  2. SQL expression (F.expr("...")) if it reads more clearly.
  3. pandas UDF for vectorisable Python logic.
  4. Plain Python UDF only when nothing else works.

Example

# Instead of a UDF that classifies amounts:
from pyspark.sql import functions as F
banded = F.when(F.col("amount") >= 1000, "large").when(F.col("amount") >= 100, "medium").otherwise("small")

Common mistakes

  1. Row-by-row API or database calls from a UDF.
  2. Forgetting None handling.
  3. Wrong declared return type producing NULLs.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Examples run on PySpark 4.2 in local mode; behaviour notes say where Spark 3.x differs

Progress is saved in this browser only. No account needed.

Search
Filter by type