Course · Data platforms
Databricks
Databricks is a lakehouse platform built around Apache Spark and Delta Lake. Learn its workspace, compute, jobs and Unity Catalog governance model.
- Lessons
- 3
- Interview questions
- 1
- Projects & case studies
- 0
- Reading time
- ~1 h
About this course
Databricks packages Spark, Delta Lake, notebooks, job orchestration and governance into one managed platform. For a data engineer, most of the work is ordinary Spark and SQL; what Databricks adds is managed compute, a job scheduler, a medallion-style lakehouse on Delta tables, and Unity Catalog for permissions and lineage.
Learn the workspace and how jobs run on compute first, then Unity Catalog. Databricks renames and regroups products often, so check current documentation for exact product names.
Your progress
Saved in this browser onlyPractise
- InterviewDatabricks interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetDatabricks Cheat SheetA quick Databricks reference: Unity Catalog names and grants, Delta table operations, medallion layers, job design and the compute choices that keep costs down.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
Resources
Cheat sheets
Related courses
- Apache SparkUnderstand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.
- PySparkPySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- Delta LakeDelta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.