Databricks courseLesson 3 of 3
Databricks course · Lesson 3 of 3
Unity Catalog and Data Governance Fundamentals
How Unity Catalog governs data in Databricks: the three-level namespace, metastores, managed versus external tables, grants, lineage and fine-grained access.
On this page
Governance means knowing what data exists, who may use it, and how it flows. In Databricks this is the job of Unity Catalog, a central catalog for tables, views, files, models and functions across workspaces.
The three-level namespace
Every table is addressed as catalog.schema.table:
SELECT * FROM prod.sales.orders;
- A catalog is the top-level container, often per environment or business domain (
dev,prod,finance). - A schema (database) groups related tables inside a catalog.
- Tables and views hold the data.
A metastore is the top-level container for catalogs, typically one per region, attached to the workspaces that use it. Because governance lives in the metastore rather than in each workspace, the same permissions apply wherever the data is accessed.
Managed and external tables
- Managed tables: Unity Catalog manages both metadata and the underlying files in a managed storage location. Dropping the table removes the data.
- External tables: metadata is in Unity Catalog, but files live in a location you manage. Dropping the table leaves the files.
- Volumes govern non-tabular files (CSVs, images, models) with the same permission model.
Permissions
Access is granted with SQL to users, groups or service principals:
GRANT USE CATALOG ON CATALOG prod TO `analysts`;
GRANT USE SCHEMA ON SCHEMA prod.sales TO `analysts`;
GRANT SELECT ON TABLE prod.sales.orders TO `analysts`;
To read a table a principal needs privileges along the path (catalog and schema usage) as well as on the table. Grant to groups, not individuals, so access follows team membership.
For finer control, Unity Catalog supports row filters and column masks, for example hiding email addresses from everyone outside a support group.
Lineage and auditing
Unity Catalog captures lineage: which jobs and notebooks read and wrote which tables and columns. That answers “what breaks if I change this column?” and “where did this number come from?”. Access is also recorded in audit logs.
A practical setup
- Catalogs per environment (
dev,staging,prod) or per domain. - Schemas per data area or medallion layer.
- Production jobs run as service principals with only the privileges they need.
- Humans get read access to
prodthrough groups; writes toprodhappen only through jobs.
Common mistakes
- Granting privileges to individual users instead of groups.
- Running production jobs with personal credentials.
- Forgetting the catalog and schema
USEprivileges and debugging “table not found” errors. - Dropping a managed table expecting the files to remain.
Interview relevance
“What is Unity Catalog used for?” and “How would you control access to sensitive columns?” are common. See the interview answer.
Key takeaway
Unity Catalog gives one namespace, one permission model and lineage across workspaces. Organise catalogs deliberately, grant to groups, run jobs as service principals, and mask sensitive columns.
Progress is saved in this browser only. No account needed.