About This Course
<div>Helios Trading Corporation needs a Data Engineer. That is you. <span style="font-size: 1rem;">Helios ships parts across six depots in the solar system. Ten source feeds land every day. Nobody has built the pipeline yet. </span><span style="font-size: 1rem;">This is not a tutorial you follow. It is a job you do. </span><span style="font-size: 1rem;">Simulate, then Execute </span><span style="font-size: 1rem;">Every build starts as a conversation. A short AI Role Play with a colleague, and goals you are measured against.</span></div><div><br></div><div>Then you build what you talked through. Most labs are work tickets: context, task, and acceptance criteria that run as real checks in the notebook.</div><div><br></div><div>What you will build</div><div><ol><li><span style="font-size: 1rem;">Design Meet Helios, walk the ten source feeds with the person who owns them, and agree the Medallion lakehouse design that every later section implements.</span></li><li><span style="font-size: 1rem;">Databricks Jobs The classic build. Auto Loader into Bronze, Structured Streaming and MERGE into Silver, SCD Type 2 for prices and customer tiers, a Gold star schema with point in time pricing, then a five task orchestrated Job with retries, alerts and a schedule.</span></li><li><span style="font-size: 1rem;">Lakeflow Spark Declarative Pipelines The same platform, rebuilt declaratively. Streaming tables, expectations as quality gates, AUTO CDC. Then an inherited pipeline fails on a night run and you diagnose it from the event log.</span></li><li><span style="font-size: 1rem;">Genie and AI/BI Dashboards A semantic layer of Metric Views over Gold, so one definition of revenue feeds both a Genie Space and a dashboard. Then you make the case for self serve to the Commercial Director.</span></li><li><span style="font-size: 1rem;">Lakebase and a Databricks App Sync curated Gold into Lakebase and deploy a live Depot Operations Console.</span></li></ol></div><div><br></div><div>Source feeds to running app. <span style="font-size: 1rem;">Real data, from a purpose built generator. </span><span style="font-size: 1rem;">The data comes from a custom synthetic generator written for this course, not a downloaded sample set. Ten interlocking feeds in three formats, around 293,000 order lines across five batches, and it is deterministic, so your numbers match the videos exactly.</span></div><div><br></div><div>The domain is order to cash for a parts distributor. Customers place orders, orders move through a lifecycle, stock leaves the shelf, prices and costs change over time, and some of it comes back as returns. Sales, inventory, pricing and returns, the same shapes you meet in retail, wholesale, manufacturing and logistics.</div><div><br></div><div>It behaves like production data too:</div><div><ul><li>Duplicate order lines</li><li><span style="font-size: 1rem;">Late arrivals from the depots with poor uplinks</span></li><li><span style="font-size: 1rem;">Dirty quantities and missing keys</span></li><li><span style="font-size: 1rem;">An order lifecycle arriving as CDC</span></li><li><span style="font-size: 1rem;">Prices that change underneath your joins</span></li></ul></div><div><span style="font-size: 1rem;">No cloud account. No credit card. No Azure subscription. </span><span style="font-size: 1rem;">The course runs end to end on Databricks Free Edition. The skills transfer directly to Azure, AWS and GCP. </span><span style="font-size: 1rem;">Every lab bootstraps to a clean state in under two minutes, so you can start at any section.</span></div><div><br></div><div><span style="font-size: 1rem;">Current for 2026</span></div><div><br></div><div>Built on what Databricks ships today:</div><div><ul><li>Lakeflow Spark Declarative Pipelines, formerly Delta Live Tables</li><li><span style="font-size: 1rem;">Lakeflow Jobs</span></li><li><span style="font-size: 1rem;">Unity Catalog and Volumes</span></li><li><span style="font-size: 1rem;">Serverless compute</span></li><li><span style="font-size: 1rem;">Genie Spaces</span></li><li><span style="font-size: 1rem;">Metric Views and AI/BI Dashboards</span></li><li><span style="font-size: 1rem;">Lakebase and Databricks Apps</span></li></ul></div><div><span style="font-size: 1rem;">These patterns also cover much of the ground tested by the Databricks Certified Data Engineer Associate and Professional exams, though this is a course about doing the job, not passing an exam.</span></div><div><br></div><div>Your instructor</div><div><br></div><div>Malvik Vaghadia is a Databricks Partner Champion and a Principal Data Engineering Consultant. He has taught more than 250,000 students across 15 courses on Databricks, PySpark, Delta Lake and Azure.</div><div><br></div><div>Helios is waiting. Start your first day.</div>
What you'll learn:
- Build a production Databricks lakehouse end to end, from ten raw source feeds to a live operational app
- Make the design decisions a data engineer is paid for, and defend them in review
- Handle source data as it really arrives: duplicates, late rows, CDC and slowly changing history
- Choose between imperative and declarative pipelines with the confidence of having shipped both
- Own a pipeline in production, keeping it trustworthy and fixing it when it fails overnight