Loading…
Lakehouse
46 posts about Lakehouse. Every summary links to the original.
Data Warehouse Modernization: Roadmap, Architecture, and Services
Data warehouse modernization addresses legacy systems that cannot scale efficiently with growing data volumes, real-time analytics, machine learning, and self-service access. The proposed roadmap spans two to four years for large estates, moving from assessment and architecture design through high-impact workload migration, governance embedding, and optimization rather than relying on a risky big-bang cutover. Its target architecture favors a lakehouse or enhanced cloud data warehouse, with open formats such as Apache Iceberg or Delta Lake, separate compute and storage, and Bronze, Silver, and Gold layers supporting incremental ELT and lineage. The source says modernization can reduce infrastructure maintenance costs by 30–50%, compress query latency from hours to seconds, and halve redundant ETL pipelines, while also improving governance for sensitive data and enabling BI, machine learning, and generative AI workloads on a shared foundation.
Databricks StaffIntroducing the Agentic CDP: A New Species of CDP for a New Era of Agents
Traditional customer data platforms (CDPs) were built for human-managed, batch-based campaigns, but the post argues that agentic buying requires millisecond speed, hyper-personalization, and richer context. It contrasts the familiar Golden Record with Golden Context, which combines customer data with current business goals and the history and outcomes of prior decisions. The proposed Agentic CDP uses “Infinity Campaigns,” always-on engagement loops that use LLMs and agents to adapt messaging, timing, and channels for individuals. It is also embedded in the data foundation, bringing customer, business, and decision context together under existing governance, and is designed for agents and humans from the outset. Databricks presents CustomerLake as an implementation of these principles for its platform.
Tasso Argyros, Ali Ghodsi, Reynold XinWhat is data pipeline architecture?
Data pipeline architecture is the end-to-end blueprint for collecting, processing, storing and delivering data from source systems to people, applications and models. It distinguishes logical design, which defines stages and responsibilities, from physical design, which assigns tools and infrastructure to those stages. The common four layers are ingestion, processing and transformation, storage, and serving and consumption, with orchestration and observability spanning the pipeline. The brief compares batch, streaming, Lambda, Kappa and medallion patterns, explaining trade-offs involving freshness, cost, complexity and operational burden. It also contrasts ETL with ELT and presents governance, monitoring and right-sized processing as reliability practices, concluding that architecture should match the use case and balance freshness, cost and reliability.
Databricks StaffLakeflow: A new era of agentic data engineering
Databricks announces a major evolution of Lakeflow, its unified platform for data engineering across ingestion, transformation, and orchestration, with capabilities centrally governed by Unity Catalog. Genie Code and generally available Lakeflow Designer support agentic and no-code pipeline development, while Genie ZeroOps monitors production assets, analyzes failures, proposes fixes, and validates them in a governed sandbox before human approval. Lakeflow Connect expands to more than 100 managed connectors, and Zerobus Ingest adds Kafka-compatible, gRPC, REST, SDK, and OpenTelemetry interfaces for high-volume event ingestion. Real-Time Mode for Spark Declarative Pipelines reaches Public Preview with end-to-end latency as low as 5 milliseconds, alongside declarative APIs and expanded Lakeflow Jobs integrations. The release also adds data-readiness triggers and external orchestration for systems including Snowflake, REST APIs, Slack, and PagerDuty.
Bilal Aslam, Ray Zhu, Manish Dalwadi, Saad Ansari, Giselle GoicocheaUnifying Data and Governance in the Agentic Era: What’s New with Azure Databricks
At Data + AI Summit 2026, Azure Databricks announced capabilities aimed at moving enterprises from experimental AI pilots to production-grade automated workflows by unifying data, productivity tools, marketing, and governance on Azure. Its Agentic Data foundation introduces LTAP, combining analytical data, streaming pipelines, and live application transactions in one lakehouse storage copy; Lakebase adds a managed serverless Postgres engine with copy-on-write branching, while Lakehouse//RT targets millisecond responses for high-concurrency workloads. Genie integrations for Microsoft Teams, M365 Copilot, Excel, and SharePoint bring governed lakehouse intelligence and ingestion into daily work, alongside tools for agents, applications, pipelines, and autonomous operations. CustomerLake adds Profile Agents and Campaign Agents for customer profiles and personalization, while Genie Ontology and Unity AI Gateway provide semantic context, rate limits, content filtering, and spend controls.
Isaac Gritz, Toussaint Webb, Ben Tripp, Kiriana StukasIntroducing CustomerLake: The Agentic CDP embedded in Databricks
Databricks announces CustomerLake, an Agentic Customer Data Platform embedded natively in its lakehouse, bringing Customer 360, identity resolution, audience building, campaign automation, activation, and personalization alongside governed data and AI models. The announcement addresses fragmented identities, stale audiences, manual campaign workflows, and the duplication and governance burden created by separate martech systems. CustomerLake uses Unity Catalog and Lakehouse Federation to access customer data across Databricks, Snowflake, Google BigQuery, cloud object storage, operational databases, and other enterprise systems, while Profile Agents create business-ready profiles and Campaign Agents build audiences, recommend actions, activate channels, and optimize engagement. Its operating model is described as embedded, democratized, and autonomous, with Agentic Identity Resolution combining deterministic, probabilistic, and agentic workflows. CustomerLake is now available in Private Preview and launches with an open partner ecosystem.
Tasso Argyros, Justin DeBrabant, Michael Trapani, Dan Morris, Katy YuanWhat is an open lakehouse? Open data standards, explained.
The piece defines an open lakehouse as a lakehouse whose storage, table format, processing engine, catalog, and ML and AI tooling use open standards and remain interchangeable. It contrasts this architecture with warehouses, lakes, and proprietary lakehouses, emphasizing low-cost object storage, ACID transactions, governance, schema guarantees, and the ability to change engines without rewriting data. Its reference stack combines open table formats such as Delta Lake and Apache Iceberg with Apache Parquet, Apache Spark, Unity Catalog, and MLflow, while allowing engines including DuckDB, Trino, and PyIceberg to work on the same data. The article also distinguishes open standards from open-source code, explains that a table format is only one layer of the stack, and states that the components can be self-hosted or consumed through a managed service.
Lisa CaoIntroducing Lakehouse//RT: Real-Time Performance on a Unified Lakehouse
Databricks introduces Lakehouse//RT, a real-time data warehouse designed for operational analytics, BI, app serving, and observability workloads. It is powered by Reyden and is intended to deliver millisecond performance directly on lakehouse data without copying it into a separate serving layer. Preview participants saw up to 16x better performance, with response times as low as 10ms on smaller datasets, sub-100ms on larger ones, and sub-100ms latency at 12,000 queries per second on standard analytical benchmarks. Tests covering concurrency, dataset scale, and complex TPCDS queries report low latency where alternatives slowed or failed. Lakehouse//RT is in Beta for select read-only workloads, with incremental autoscaling and automatic baseline compute sizing.
Nong Li, Shoumik Palkar, Shant Hovsepian, Mostafa Mokhtar, Reynold XinWhat is enterprise intelligence?
Enterprise intelligence (EI) is presented as an organization-wide capability combining business intelligence, knowledge management, enterprise search and AI to turn structured and unstructured data into decisions and actions. Unlike traditional BI, which centers on dashboards, reports and structured data, EI connects these capabilities through a shared architecture and governed business context. The described stack includes a lakehouse-based data foundation, batch and streaming pipelines, governance, semantics, analytics, search, machine learning, generative AI and a decision layer. Shared definitions such as “active customer” and “monthly revenue” are intended to keep dashboards, queries and AI agents aligned, while maintained context addresses knowledge that becomes stale as the business changes. The result described is a common trusted source from which people, applications and agents can produce insights and initiate actions.
Databricks StaffTalk to all your data, wherever it lives
Lakehouse Federation addresses the challenge of reasoning across enterprise data spread among AWS Glue, Snowflake, Oracle, BigQuery, Postgres, and legacy formats without first migrating it. By connecting external sources in place and syncing their metadata into Unity Catalog, Databricks applies shared permissions, lineage, and access controls while preserving source data. Federated comments and descriptions give Genie schema context, while Unity Catalog Semantics lets teams define governed metrics such as ROI once for consistent use across Genie, dashboards, and notebooks. The example connects an AWS Glue marketing database, carries its metadata, defines an ROI metric view, and asks Genie which campaigns led ROI last quarter. The post reports an immediate, accurate answer from live Glue data, and describes planned richer semantics, broader federation, and possible performance gains from managed tables.
John SpencerUnlocking semantics for AI: How Mercedes-Benz Korea built trusted “Talk to Data” at scale
Mercedes-Benz Korea piloted a “Talk to Data” architecture that extends its Databricks analytics foundation with a governed semantic layer for enterprise AI, rather than treating the effort as a chatbot project. The design moves Power BI DAX KPI logic into Unity Catalog Business Semantics and Metric Views, keeping sources, joins, measures, dimensions, comments, and synonyms alongside governed Lakehouse data. Genie spaces use curated metric views for domain questions, while Agent Bricks composes persona-based agents, with Unity Catalog enforcing row- and column-level access. An automated DAX-to-Metric-View transpiler parses semantic models, maps tables, generates draft definitions, flags non-automatable measures, and reports conversion gaps. The documented playbook combines gold-layer curation, KPI validation, regression testing, Genie optimization, persona agents, and Databricks Apps; the pilot reports AI answers aligned with established KPI definitions and BI reporting logic.
Sai Yang, Fares Kamal, Alina Kamal, Andreas Jäck, Johannes Laufer, Manuel CulebrasForward Deployed Engineering: Delivering Business Outcomes with AI
Databricks is formalizing its Forward Deployed Engineering (FDE) organization to address customers’ shift from migration and data-pipeline requests toward business outcomes with AI. FDE brings Professional Services together around an engineering-led model that embeds engineers with customers, supports modernization and production AI, and works with partners and Databricks R&D. In the cited examples, teams migrated five-plus petabytes of JPMC Consumer and Community Banking Risk data and more than 500 notebooks in four months, while Fox used Lakebase, AI Search, Databricks Apps, and Model Serving to redesign fan experiences. The organization says its engagements use shared OKRs, rapid prototype-to-production delivery, embedded engineering, outcome-aligned commercial options, and global partner coverage, with Fox reporting that Sports AI users spend approximately twice as long in the app.
Jason MartinHow Ecolab rebuilt retail intelligence on Databricks and Anthropic Claude
Ecolab needed to combine audits, health inspections, pest telemetry, and other data from nine systems so retail teams could answer location-specific compliance questions. Its Retail Intelligence application is a native Databricks App using Lakebase Postgres, Lakeflow, and Spark Declarative Pipelines to move governed data into a Unity Catalog lakehouse, while Foundation Model APIs serve Claude Sonnet, Claude Haiku, and Gemini. A Coordinator Agent delegates requests to specialized agents that use Vector Search, SQL, Unity Catalog Functions, and an external MCP server; a Response Agent returns cited answers, with short- and long-term memory stored through Lakebase. The system also applies five Judge LLMs, MLflow tracing, and ai_query() batch inference. Report preparation fell from two weeks to under two minutes, while the assistant supports approximately twelve languages at about 98% accuracy.
Babu Chinnaswamy, Nicholas Dylla, Alissa Ellingson, Harish GaurScaling AI Through Data Fluency
Aer Lingus is redirecting a significant share of its IT and change spending from traditional maintenance toward a Databricks-powered data foundation, addressing legacy systems that trap information in departmental silos. Dave O’Donovan says the airline spent the past 18 months prioritizing platform development, governance, data quality and data literacy rather than chasing each new AI product. Databricks was selected for a unified lakehouse architecture, with data warehousing, Genie’s plain-English querying and real-time operational data intended to broaden access beyond specialist teams. At Aer Lingus’s Operations Control Center, combining sensor and operational inputs gives teams a fuller real-time view for disruption decisions, while commercial teams use live insights to adjust pricing. The transformation also includes a Data Literacy Academy, a 75/25 capacity split between foundational work and innovation, a 20-person Continuous Improvement team, and experiments with agents for business-case development and CFO review.
Aly McGueJumpstart your Data Modeling with Databricks Industry Data Models
Databricks is publishing a public library of 40 Lakehouse industry data models designed to provide Silver-layer foundations for analytics and machine learning. Each industry offers a Minimum Viable Model and Expanded Coverage Model derived from the same model.json, with breadth rather than attribute depth distinguishing the scopes. A rules-driven AI agent applies more than 200 structural checks across 14-plus modeling domains, enforcing hierarchy, primary and foreign keys, normalization, division balance, data types, governance tags, and acyclic relationships. The models deploy to Unity Catalog in three physical cataloging styles and include DDL, schemas, metric views, classification tags, ontology, diagrams, and synthetic data with valid references. The airline ECM example contains 19 domains, 420 products, 17,278 attributes, 420 primary keys, and 2,877 foreign keys, while the models remain customizable starting points requiring domain expertise and organizational review.
Amr Ali, Drew Triplett, Franco Patano, Shelley ShafferyModern BSA/AML compliance on Databricks
AML operations are strained by fragmented systems, high false-positive volumes, manual case documentation, and opaque vendor scoring, leaving analysts focused on backlog rather than financial-crime intelligence. The proposed Databricks Data + AI Platform unifies transaction monitoring, KYC, sanctions, case history, and policy data under Unity Catalog, using Lakeflow Connect and a Bronze–Silver–Gold Delta architecture with masking, row-level security, and lineage. MLflow, Model Serving, Lakehouse Monitoring, and inference tables support institution-specific detection models, while Agent Bricks coordinates agents for evidence gathering, recommendations, and SAR drafting with analysts retaining final decisions. The architecture also uses Lakebase for governed operational state and Databricks Apps for analyst and executive experiences. Reported outcomes include a 75% reduction in false positives reaching the analyst queue and compressing three-to-six-hour investigations to minutes.
Kateryna Savchyn, Pavithra Rao, Mimi Park, Emerson BayukAnnouncing the winners of the 2026 Databricks Customer Awards
The 2026 Databricks Customer Awards recognize organizations and leaders using the Databricks Data + AI Platform across eight categories and four regions. The announcement names winners including Applied Materials, Virgin Atlantic, Fonterra Co-operative Group, Telefónica | Vivo, Virtue Foundation, Octopus Energy, Axpo, Atlassian, Wassym Bensaid at Rivian and Volkswagen Group Technologies, and Kenan Colson at Lippert. Examples include Applied Materials’ move from a Hadoop-based data lake to a governed lakehouse, with 1,500-plus analysts, more than 100 production machine-learning models and faster pipeline development, while Fonterra centralized data to improve supply-chain and compliance work. Rivian unified vehicle, factory and enterprise data, consolidated legacy systems into Delta and Unity Catalog, and designed for 500–600 petabytes; Lippert deployed AI tools across customer care, finance and HR. The announcement presents data and AI adoption as a cross-functional operating model.
Sara SteffenAnnouncing the 2026 Databricks Customer Awards Industry winners
Databricks announced its 2026 Customer Awards Industry winners, recognizing 10 organizations across financial services, communications, health and life sciences, manufacturing, retail and CPG, energy and utilities, enterprise technology, public sector, digital-native businesses and cybersecurity. The cited work uses data and AI to address industry-specific needs, including SMBC Group’s governed lakehouse for risk and finance, Hospital for Special Surgery’s full-system ingestion strategy, Lumen’s conversational service-operations workflows and Superhuman’s high-volume model serving. Reported results include HSS ingesting more than 40 source systems and creating over 14,500 production tables, Lumen recording 3 million-plus AI-powered diagnostics and 35% ticket deflection, and Superhuman handling peaks above 200,000 queries per second. Adobe is also recognized for applying software engineering practices to cybersecurity detection workflows, reducing false positives and improving development speed.
Michael GriffithsPractical Data Lakehouse Examples and Use Cases
The article presents practical data lakehouse patterns for streaming analytics, IoT pipelines, machine learning workflows, and enterprise reporting, addressing the gap between theoretical definitions and deployable examples. It explains how a lakehouse combines low-cost, schema-flexible object storage with schema enforcement, ACID transactions, data versioning, lineage tracking, and query performance, allowing SQL and ML workloads to use shared open-format tables. Examples include second-level fraud detection, medallion-based historical analytics, predictive maintenance from sensor data, and governed customer 360 profiles. The implementation guidance covers raw storage and partitioning, centralized catalogs, decoupled compute, role-based access, time travel, migration coexistence, SLAs, observability, and lifecycle policies. Together, these patterns are presented as a unified architecture that reduces duplication and data movement while supporting governed analytics at scale.
Databricks StaffData Governance Architecture: A Complete Blueprint for Modern Organizations
Data governance architecture is presented as a blueprint for aligning policies, roles, processes, and technologies with business outcomes. It defines objectives including consistent data definitions, data integrity, layered access controls, and secure self-service analytics, while assigning responsibilities across executives, architects, engineers, analysts, managers, and compliance teams. Core principles are accountability, transparency, consistency, and stewardship, supported by federated ownership through councils, data owners, and embedded stewards. The discussion compares DAMA-DMBOK, TOGAF, and Zachman according to organizational scale, regulatory context, and architecture maturity, and describes modern patterns including lakehouse, data mesh, and data fabric. It concludes that effective programs require executive sponsorship, documented roles, measurable quality controls, iterative implementation, and sustained change management.
Databricks Staff