Menu

Course · Data platforms

Databricks

Databricks is a lakehouse platform built around Apache Spark and Delta Lake. Learn its workspace, compute, jobs and Unity Catalog governance model.

Lessons
3
Interview questions
1
Projects & case studies
0
Reading time
~1 h

About this course

Databricks packages Spark, Delta Lake, notebooks, job orchestration and governance into one managed platform. For a data engineer, most of the work is ordinary Spark and SQL; what Databricks adds is managed compute, a job scheduler, a medallion-style lakehouse on Delta tables, and Unity Catalog for permissions and lineage.

Learn the workspace and how jobs run on compute first, then Unity Catalog. Databricks renames and regroups products often, so check current documentation for exact product names.

Your progress

Saved in this browser only

Practise

Course structure

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. Databricks for Data EngineersWhat a Data Engineer needs to know about Databricks: Spark and Delta Lake underneath, workspaces, compute, jobs, medallion pipelines, Unity Catalog and cost control.Intermediate2 min

Beginner

Core concepts you will use every day.

  1. Databricks Workspace, Jobs and Lakehouse ConceptsGet oriented in Databricks: workspaces, notebooks and Git folders, all-purpose versus job compute, scheduled jobs, and the medallion lakehouse on Delta tables.Beginner2 min

Intermediate

Patterns used in production pipelines.

  1. Unity Catalog and Data Governance FundamentalsHow Unity Catalog governs data in Databricks: the three-level namespace, metastores, managed versus external tables, grants, lineage and fine-grained access.Intermediate2 min

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type