Course · Languages & query
Python
Python glues pipelines together: ingestion, validation, orchestration and PySpark jobs. Focus on functions, generators, error handling and testable code.
- Lessons
- 5
- Interview questions
- 4
- Projects & case studies
- 3
- Reading time
- ~1 h
About this course
In Data Engineering, Python is mostly about reliable plumbing: reading from sources, validating records, loading targets, scheduling work and calling libraries such as PySpark. The difference between a script and a pipeline is error handling, logging and the ability to run the same job twice safely.
Start with functions and modules, then iterators and generators, then exceptions and logging.
Your progress
Saved in this browser onlyPractise
- InterviewPython interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetPython for Data Engineers Cheat SheetA quick Python reference for pipelines: files and CSV, JSON, dates, collections, generators and batching, error handling, logging and testing with pytest.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
- Python Functions, Modules and Reusable Data Pipeline CodeStructure Python pipeline code as small pure functions, clear modules and a thin entry point so it is testable, reusable and safe to run from any scheduler.
- Python Data Structures for Data Engineering InterviewsChoose between list, tuple, set, dict, deque and Counter by their operations and costs, with the patterns that come up in data engineering coding interviews.
Intermediate
Patterns used in production pipelines.
- Python Iterators, Generators and Memory-Efficient ProcessingUse Python iterators and generators to process files and API pages lazily, in batches, with flat memory use, and avoid the mistakes that silently exhaust them.
- Build an Idempotent CSV Loader in PythonBuild a small Python loader that reads CSV files, skips and counts bad rows, and loads into a database so running it twice gives the same result.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- BeginnerCSV to Data Warehouse PipelineA small online shop exports orders as daily CSV files. Build a pipeline that loads them into a star schema so that sales can be reported reliably, even when files are resent or contain bad rows.
- IntermediateE-commerce Analytics Data PlatformAn online store has orders, customers, products and web sessions in separate systems. Build an ELT platform that models them into trusted marts for revenue, retention and product performance.
- AdvancedFraud Detection Data PipelineBuild the data side of a fraud-detection system: compute per-card behavioural features from a transaction stream, flag suspicious transactions with transparent rules, and maintain a feature table that a model could use.
Resources
Cheat sheets
Related courses
- SQLSQL is the core language of data work: querying, transforming and modelling data in warehouses, lakehouses and Spark. Start here before any other tool.
- PySparkPySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- AirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.