Loading…
How a major freight railroad scaled pipeline creation with Genie Code
Dinesh Chandrasekaran, Subhadip Chanda, Julia Powell, Gal Oshri
- Source
- Databricks
- Published
- Added to Yomu
Summary
A major Canadian freight railroad with roughly 20,000 route miles needed to modernize a decades-old analytical estate while handling hundreds of pipelines and rising demand for real-time analytics and AI. Using Databricks Genie Code, Unity Catalog, custom Agent Skills and a Streamlit Databricks App, the team created a metadata-grounded pipeline-generation workflow driven by compact YAML prompts. The workflow discovers schemas, maps source fields, and emits six artifacts—DDL, historical and streaming ingestion, incremental merges, and automated tests—using deterministic patterns for audit columns, deduplication, change-sequence guards and soft-delete reconciliation. Human-reviewed Source-to-Target Mapping preserves business interpretation, while Genie Code handles discovery, orchestration and artifact generation within governed Databricks execution. The reported result was more than 90% automation for new table ingestion, reducing delivery from days per table to minutes and supporting single, multi-table and bulk generation.
Context
The company was modernizing a decades-old analytical estate spanning mainframes, legacy data warehouses, enterprise ETL platforms and appliances. Building one pipeline required multiple days of schema inspection, business-logic translation, ingestion and merge development, downstream transformations, and testing. That manual effort could not scale to hundreds of tables while preserving critical legacy logic and supporting real-time analytics and AI.
Approach / What changed
The team combined Genie Code, Unity Catalog, custom versioned Agent Skills and a Streamlit app built on Databricks Apps. Data designers review Source-to-Target Mappings, while compact YAML prompts drive metadata discovery and deterministic generation of DDL, historical loads, streaming ingestion, incremental merges and automated tests. Explicit templates and invariants enforce enterprise conventions for audit columns, deduplication, change-sequence guards, type casts and soft-delete reconciliation.
Takeaways
- A custom Agent Skill packages ingestion standards, naming conventions and artifact patterns so Genie Code can apply them consistently across generated pipelines.
- Unity Catalog provides live schema introspection across raw, historical and prep layers, grounding column matching, transformation inference and validation in current metadata.
- The workflow reports more than 90% automation for new table ingestion and reduces pipeline delivery from days per table to minutes, with single-table, multi-table and bulk modes.