Databricks Data Engineering: the lakehouse, engineered.
Four days on the Databricks platform the way production teams actually run it: Delta Lake internals, Spark done right, declarative pipelines with DLT, Unity Catalog governance, streaming, and performance tuning.
What this course is
A practitioner course for teams adopting or scaling Databricks — taught by engineers who run lakehouses in production. We go past the notebooks-and-demos level into the internals that decide whether your platform is fast, governed, and affordable.
Who should attend
- Data engineers building or migrating to Databricks
- Analytics engineers owning dbt/Databricks SQL workloads
- Architects designing lakehouse governance and cost models
What you need coming in
- Professional SQL; basic Python or Scala readability
- Familiarity with data warehousing concepts (facts, dimensions, SCDs)
- Databricks workspaces provided for labs; no setup required
Learning objectives
- Explain Delta Lake internals — transaction log, OPTIMIZE, Z-ORDER, liquid clustering — and apply them to real tables
- Write efficient Spark: partitioning, join strategies, shuffle control, and Photon-aware patterns
- Build declarative pipelines with Delta Live Tables including expectations (data quality gates)
- Implement Unity Catalog: catalogs, schemas, grants, row/column security, and lineage
- Ingest streaming data with Auto Loader and Structured Streaming, with exactly-once semantics
- Diagnose slow jobs from the Spark UI and query history; tune for cost and latency
Day by day
Day 1 — Lakehouse foundations & Delta Lake internals
Medallion architecture done right. The Delta transaction log, time travel, schema enforcement vs. evolution, deletes and GDPR patterns, OPTIMIZE, Z-ORDER, and liquid clustering. Storage layout decisions that compound.
LABS → build bronze/silver/gold tables; break and fix schema evolution; time-travel a bad load; benchmark Z-ORDER on a skewed datasetDay 2 — Spark engineering & Delta Live Tables
Spark execution: stages, shuffles, skew, and join strategies. Reading the Spark UI like a profiler. Then DLT: declarative pipelines, expectations as data-quality gates, and pipeline observability.
LABS → fix a skewed join two ways and measure; convert a notebook pipeline to DLT with expectations; set up pipeline alertsDay 3 — Unity Catalog governance
Metastore design, catalog/schema/table grants, service principals and groups, row filters and column masks, lineage and audit. The governance model auditors accept and engineers tolerate.
LABS → stand up a governed catalog; implement PII masking; trace lineage end to end; run an access-review drillDay 4 — Streaming, performance & capstone
Auto Loader, Structured Streaming checkpoints and watermarks, exactly-once sinks. Performance clinic: query history analysis, Photon, caching, and cost attribution. Capstone: a complete medallion pipeline with governance and quality gates.
LABS → stream files with Auto Loader; tune a slow job from the Spark UI; capstone build + reviewDelivery options
4-day course
The complete syllabus, onsite or virtual, up to 20 participants. Lab workspaces provided. Includes a post-course office-hours session for follow-up questions.
2-day essentials
Days 1–2 condensed: lakehouse foundations, Delta internals, and DLT. For teams that need builders productive immediately.
1-day exec workshop
Lakehouse strategy, Databricks cost models, governance, and build-vs-buy for leadership. Available on request.
Private cohorts only, tailored to your cloud (AWS, Azure, GCP) in a pre-course scoping call. Migrating to Databricks? See our database modernization guides and the Academy.
Make your team dangerous on Databricks.
Tell us your team size, cloud, and timeline — we'll scope the cohort and send a proposal.