Data Engineer Track
Platform depth for the Databricks Data Engineer interview — ten modules in reading order, from Spark fundamentals to the certification roadmap. Each module ends in interview questions and story hooks from real production work.
Apache Spark Internals & Troubleshooting
Architecture, lazy evaluation, Catalyst & AQE, shuffles, skew, spill, and how to debug all of it in the Spark UI.
MODULE 02Delta Lake Deep Dive
The transaction log, ACID on object storage, OPTIMIZE, Z-ORDER vs. liquid clustering, deletion vectors, time travel, and Change Data Feed.
MODULE 03Streaming & Incremental Ingestion
Structured Streaming, Auto Loader, COPY INTO, triggers, watermarks, exactly-once semantics, and backfill patterns.
MODULE 04Lakeflow: Jobs & Declarative Pipelines
Orchestration on the platform — Lakeflow Jobs (Workflows), Lakeflow Declarative Pipelines (formerly DLT), expectations, and how to choose between them.
MODULE 05Unity Catalog & Governance
The UC object model, privileges, lineage, row filters and column masks, system tables, and Delta Sharing.
MODULE 06Lakehouse Data Modeling
Medallion architecture done right, dimensional modeling on Delta, CDC, SCD Type 1/2 patterns, and data-quality design.
MODULE 07Performance Tuning & FinOps
Cluster sizing, Photon, file-size management, caching, join strategies, and the cost observability that turns tuning into dollars.
MODULE 08Databricks SQL & the Semantic Layer
SQL warehouses, Photon, AI/BI dashboards and Genie, and integrating Power BI and external semantic layers.
MODULE 09DevOps on Databricks
Databricks Asset Bundles, CI/CD pipelines, testing PySpark, environment promotion, and Terraform.
MODULE 10Certification Roadmap
Databricks Data Engineer Associate and Professional exam blueprints, mapped against existing Microsoft credentials and real experience.