Databricks courseLesson 2 of 3
Databricks course · Lesson 2 of 3
Databricks Workspace, Jobs and Lakehouse Concepts
Get oriented in Databricks: workspaces, notebooks and Git folders, all-purpose versus job compute, scheduled jobs, and the medallion lakehouse on Delta tables.
On this page
Databricks is a managed platform for Spark, SQL and machine learning on a lakehouse. The concepts below are the ones a data engineer uses every day.
Workspace
The workspace is where you organise and run work: notebooks, files, SQL queries, dashboards, and Git folders that connect a folder to a Git repository. Treat Git as the source of truth for production code; notebooks are fine for exploration, but production jobs should run code that is version-controlled and reviewed.
Compute
| Type | Use |
|---|---|
| All-purpose (interactive) compute | Development and exploration, shared by people |
| Job compute | Created for a scheduled job run and shut down after; usually cheaper for production |
| SQL warehouses | SQL and BI workloads |
| Serverless options | Compute managed by Databricks, with fast start-up |
Run production pipelines on job compute (or serverless job compute), not on an interactive cluster someone might restart.
Jobs
A job runs one or more tasks (a notebook, a Python script or wheel, a SQL query, a pipeline) on a schedule or trigger, with dependencies, retries, parameters, notifications and run history. The same design rules as any orchestrator apply: make each task idempotent, pass the processing date as a parameter, and alert on failure. Teams that already use Airflow can trigger Databricks jobs from Airflow instead.
The medallion lakehouse
Databricks popularised organising Delta tables in layers:
| Layer | Contents | Typical rules |
|---|---|---|
| Bronze | Raw data as ingested | Append, keep source fidelity, add ingestion metadata |
| Silver | Cleaned, deduplicated, conformed | Typed schemas, quality checks, merges |
| Gold | Business-level aggregates and models | Star schemas and marts for BI and ML |
Each layer is a set of Delta tables. Because Delta provides ACID commits and time travel, each step can be rerun safely and audited.
Where Spark fits
Most Databricks pipeline code is ordinary PySpark or SQL. Everything in the Spark execution model and partitions and skew guides applies unchanged.
Common mistakes
- Running production jobs on shared interactive clusters.
- Keeping production logic only in notebooks without version control or tests.
- Skipping the silver layer and building gold tables directly on raw data.
- Leaving interactive clusters running without auto-termination.
Interview relevance
Interviewers ask how you would structure a Databricks pipeline: medallion layers, job compute, idempotent tasks, Delta tables and Unity Catalog permissions.
Key takeaway
Develop in the workspace, version code in Git, run production on job compute, and organise Delta tables into bronze, silver and gold layers.
Progress is saved in this browser only. No account needed.