Track 01 · Data Engineer

Data Engineer Track

Platform depth for the Databricks Data Engineer interview — ten modules in reading order, from Spark fundamentals to the certification roadmap. Each module ends in interview questions and story hooks from real production work.

MODULE 01

Apache Spark Internals & Troubleshooting

Architecture, lazy evaluation, Catalyst & AQE, shuffles, skew, spill, and how to debug all of it in the Spark UI.

MODULE 02

Delta Lake Deep Dive

The transaction log, ACID on object storage, OPTIMIZE, Z-ORDER vs. liquid clustering, deletion vectors, time travel, and Change Data Feed.

MODULE 03

Streaming & Incremental Ingestion

Structured Streaming, Auto Loader, COPY INTO, triggers, watermarks, exactly-once semantics, and backfill patterns.

MODULE 04

Lakeflow: Jobs & Declarative Pipelines

Orchestration on the platform — Lakeflow Jobs (Workflows), Lakeflow Declarative Pipelines (formerly DLT), expectations, and how to choose between them.

MODULE 05

Unity Catalog & Governance

The UC object model, privileges, lineage, row filters and column masks, system tables, and Delta Sharing.

MODULE 06

Lakehouse Data Modeling

Medallion architecture done right, dimensional modeling on Delta, CDC, SCD Type 1/2 patterns, and data-quality design.

MODULE 07

Performance Tuning & FinOps

Cluster sizing, Photon, file-size management, caching, join strategies, and the cost observability that turns tuning into dollars.

MODULE 08

Databricks SQL & the Semantic Layer

SQL warehouses, Photon, AI/BI dashboards and Genie, and integrating Power BI and external semantic layers.

MODULE 09

DevOps on Databricks

Databricks Asset Bundles, CI/CD pipelines, testing PySpark, environment promotion, and Terraform.

MODULE 10

Certification Roadmap

Databricks Data Engineer Associate and Professional exam blueprints, mapped against existing Microsoft credentials and real experience.