PySpark interview questionsQuestion 1 of 2
PySpark interview question · Question 1 of 2
When should you avoid Python UDFs in PySpark?
Short answer
Avoid a Python UDF whenever built-in Spark functions can express the logic. A UDF moves data between the JVM and Python workers and is a black box to the optimiser, so it is slower and prevents optimisations such as predicate pushdown. If Python is genuinely needed, prefer a vectorised pandas UDF that processes batches through Apache Arrow, and keep any UDF pure, null-safe and tested.
Detailed explanation
Built-in functions run in the JVM with code generation and are understood by the optimiser. A Python UDF adds a separate Python evaluation step in the physical plan. In Spark 4, regular Python UDFs can use Arrow for transfer by default when PyArrow is installed, which narrows the serialisation gap, but the function is still opaque to the optimiser and still executes Python.
Decision order
- Built-in function (
when,regexp_extract,date_trunc, array functions). - SQL expression (
F.expr("...")) if it reads more clearly. - pandas UDF for vectorisable Python logic.
- Plain Python UDF only when nothing else works.
Example
# Instead of a UDF that classifies amounts:
from pyspark.sql import functions as F
banded = F.when(F.col("amount") >= 1000, "large").when(F.col("amount") >= 100, "medium").otherwise("small")
Common mistakes
- Row-by-row API or database calls from a UDF.
- Forgetting
Nonehandling. - Wrong declared return type producing
NULLs.
Progress is saved in this browser only. No account needed.