---
title: "Databricks"
description: "Data and AI platform for data engineering, analytics, machine learning, and generative AI."
---

# Databricks
> Data and AI platform for data engineering, analytics, machine learning, and generative AI.

## Articles

### [Accelerate search queries with full-text search indexes on Databricks](https://yomu.fyi/post/accelerate-search-queries-with-full-text-search-indexes-on-databricks.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Yu Xu, Yingyi Bu, Ivan Vezilić
- Published: Jun 16, 2026

Databricks introduces full-text search indexes in Beta on Databricks Runtime 18.2 to accelerate substring and keyword queries on large open-format tables without changing their layouts. The indexes tokenize text columns into a compact lookup structure mapping tokens to matching rows; at query time, the engine uses it to identify candidate files and skip most of the table. They are maintained asynchronously, require no query hints, preserve complete results when stale by scanning indexed and non-indexed data as needed, and support Unity Catalog managed Delta and Iceberg tables on serverless and classic compute. A Trust and Safety team reported a substring search running more than 100x faster on a petabyte-scale table, while Liquid clustering remains complementary because it optimizes column-value filters rather than text within fields.


### [Introducing CustomerLake: The Agentic CDP embedded in Databricks](https://yomu.fyi/post/introducing-customerlake-the-agentic-cdp-embedded-in-databricks.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Tasso Argyros, Justin DeBrabant, Michael Trapani, Dan Morris, Katy Yuan
- Published: Jun 16, 2026

Databricks announces CustomerLake, an Agentic Customer Data Platform embedded natively in its lakehouse, bringing Customer 360, identity resolution, audience building, campaign automation, activation, and personalization alongside governed data and AI models. The announcement addresses fragmented identities, stale audiences, manual campaign workflows, and the duplication and governance burden created by separate martech systems. CustomerLake uses Unity Catalog and Lakehouse Federation to access customer data across Databricks, Snowflake, Google BigQuery, cloud object storage, operational databases, and other enterprise systems, while Profile Agents create business-ready profiles and Campaign Agents build audiences, recommend actions, activate channels, and optimize engagement. Its operating model is described as embedded, democratized, and autonomous, with Agentic Identity Resolution combining deterministic, probabilistic, and agentic workflows. CustomerLake is now available in Private Preview and launches with an open partner ecosystem.


### [What is customer segmentation?](https://yomu.fyi/post/what-is-customer-segmentation.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 16, 2026

Customer segmentation divides an existing customer base into smaller groups based on shared demographic, geographic, psychographic, behavioral, firmographic or value-based characteristics. Unlike market segmentation, it uses first-party data about existing customers, and a customer can belong to multiple segments at once. The guide distinguishes rule-based, survey-based, RFM, k-means, decision-tree and AI/ML-driven methods, with choices depending on data maturity and business goals. Effective implementation defines an objective, audits and unifies sources, resolves duplicate identities, selects a method, validates segments and measures outcomes. It also describes Databricks' CustomerLake capabilities, including governed Customer 360, agentic identity resolution, natural-language segmentation through Genie and bidirectional activation connectors; stated benefits include improved retention, conversion, customer lifetime value and marketing efficiency.


### [What is an open lakehouse? Open data standards, explained.](https://yomu.fyi/post/what-is-an-open-lakehouse-open-data-standards-explained.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Lisa Cao
- Published: Jun 16, 2026

The piece defines an open lakehouse as a lakehouse whose storage, table format, processing engine, catalog, and ML and AI tooling use open standards and remain interchangeable. It contrasts this architecture with warehouses, lakes, and proprietary lakehouses, emphasizing low-cost object storage, ACID transactions, governance, schema guarantees, and the ability to change engines without rewriting data. Its reference stack combines open table formats such as Delta Lake and Apache Iceberg with Apache Parquet, Apache Spark, Unity Catalog, and MLflow, while allowing engines including DuckDB, Trino, and PyIceberg to work on the same data. The article also distinguishes open standards from open-source code, explains that a table format is only one layer of the stack, and states that the components can be self-hosted or consumed through a managed service.


### [Introducing Lakehouse//RT: Real-Time Performance on a Unified Lakehouse](https://yomu.fyi/post/introducing-lakehouse-rt-real-time-performance-on-a-unified-lakehouse.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Nong Li, Shoumik Palkar, Shant Hovsepian, Mostafa Mokhtar, Reynold Xin
- Published: Jun 16, 2026

Databricks introduces Lakehouse//RT, a real-time data warehouse designed for operational analytics, BI, app serving, and observability workloads. It is powered by Reyden and is intended to deliver millisecond performance directly on lakehouse data without copying it into a separate serving layer. Preview participants saw up to 16x better performance, with response times as low as 10ms on smaller datasets, sub-100ms on larger ones, and sub-100ms latency at 12,000 queries per second on standard analytical benchmarks. Tests covering concurrency, dataset scale, and complex TPCDS queries report low latency where alternatives slowed or failed. Lakehouse//RT is in Beta for select read-only workloads, with incremental autoscaling and automatic baseline compute sizing.


### [Databricks announces 2026 global partner awards](https://yomu.fyi/post/databricks-announces-2026-global-partner-awards.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Kori O'Brien, Stephen Orban
- Published: Jun 15, 2026

Databricks announced its 2026 Partner Awards at Data + AI Summit, recognizing more than 65 achievements across its global partner network of over 8,000 organizations. The awards cover consulting and system integrators, independent software vendors, technical champions, learning and enablement, industry categories, and product-focused contributions. Accenture/Avanade received Global Partner of the Year for an eighth consecutive year, citing more than 15,000 trained practitioners, 9,500+ certified resources, and over 1,000 joint engagements. Other cited results include Capgemini’s Unity Catalog migration of 25,000 tables, 10,000 notebooks, and 2 petabytes across 150+ countries, while Kraken Technologies reduced data-processing costs eightfold and cut load times from three days to eight hours using Databricks and Delta Sharing. The announcement frames the winners’ work as supporting data and AI adoption through software, services, integrations, and consulting.


### [Announcing New OpenSharing and Marketplace capabilities for the AI era](https://yomu.fyi/post/announcing-new-opensharing-and-marketplace-capabilities-for-the-ai-era.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Akram Chetibi, Harish Gaur, Huey Han, Tia Chang, Mengxi Chen, Lin Zhou, DJ Sharkey
- Published: Jun 15, 2026

Databricks announces new OpenSharing and Marketplace capabilities aimed at sharing data, AI assets, and partner applications without moving data. OpenSharing, a Linux Foundation project and evolution of Delta Sharing, adds vendor-neutral sharing for Agent Skills, AI models, unstructured data, Iceberg clients, Lakebase tables and change data feed, plus governed multi-cloud connectivity and agentic sharing through Genie Agent Sharing. The release also introduces identity resolution in Databricks Clean Rooms, where partners such as LiveRamp and Acxiom can work against protected first-party data without seeing raw customer records. Third-party apps are now available through Databricks Marketplace, allowing customers to deploy them in their workspaces without separate infrastructure, while providers gain distribution. Marketplace Commit Drawdown lets customers use existing Databricks universal commits to acquire Marketplace data and AI solutions.


### [Skip the learning curve: rethinking data migration for real outcomes](https://yomu.fyi/post/skip-the-learning-curve-rethinking-data-migration-for-real-outcomes.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Vijay Anala
- Published: Jun 15, 2026

Data migrations are presented as high-risk, costly initiatives whose technical completion can delay adoption and strategic value, especially during infrequent warehouse transitions. The proposed alternative combines migration, modernization, and value creation in parallel, using experienced specialized partners and AI-enabled automation for code conversion, data-quality validation, and pipeline modernization. Rather than lifting and shifting legacy workloads, teams are urged to simplify architectures, retire unnecessary components, reduce technical debt, and align data with business needs while validating progress continuously. Progressive decommissioning reduces the “double-bubble” period in which old and new systems run together, helping costs fall as workloads move instead of waiting for final completion. The Migrate & Modernize Program connects organizations with partners, and the post reports faster cutovers, reduced migration costs, and complex workloads entering production ahead of schedule among early participants.


### [What is Document AI?](https://yomu.fyi/post/what-is-document-ai.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 15, 2026

Document AI uses machine learning, natural language processing (NLP) and optical character recognition (OCR) to extract, classify and understand information from structured, semi-structured and unstructured documents. Unlike OCR alone, it interprets layout and context, turning files such as invoices, contracts and emails into structured, actionable data through ingestion, OCR, layout parsing, entity extraction, classification, validation and, when needed, human review. Modern systems add large language models for summarization, document Q&A and zero-shot extraction, but hallucination risk makes validation and human oversight essential, particularly in regulated settings. The guide also describes Databricks Document Intelligence, which processes and stores documents alongside organizational data under Unity Catalog, using AI Functions, Variant and Lakeflow Jobs to create governed, queryable workflows without moving data between systems.


### [Introducing Omnigent: A Meta-Harness to Combine, Control and Share Your Agents](https://yomu.fyi/post/introducing-omnigent-a-meta-harness-to-combine-control-and-share-your.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Matei Zaharia, Kasey Uhlenhuth, Corey Zumar
- Published: Jun 13, 2026

Omnigent is an open-source meta-harness designed to combine, control, and share agents across different harnesses, models, and interfaces. Databricks says users currently juggle multiple agents and copy context between them, while builders struggle to combine or replace harnesses with incompatible interfaces. Its runner wraps terminal-based agents and SDKs in sandboxed sessions with a uniform API for messages, files, streamed text, and tool calls; a server adds policies, sharing, and access through terminal, web, mobile, native Mac OS, and APIs. Features include live collaboration, hosted sandbox execution, contextual security and cost policies, OS isolation, and multi-harness authoring. Released in alpha under Apache 2.0, Omnigent is intended to provide a durable layer above changing agents and harnesses.


### [From Wall Street to Data Platforms](https://yomu.fyi/post/from-wall-street-to-data-platforms.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Andrea Fernández, Kim Hatton
- Published: Jun 13, 2026

Kim Hatton, Databricks’ Global Financial Services Marketing Leader, describes how two decades in regulated financial-services marketing led her toward technology and data-centered strategy. She says financial institutions still need to unlock value while navigating compliance, but their divisions and systems make a unified customer view difficult. Her account points to Unity Catalog’s unified governance and single source of truth for breaking down silos and supporting requirements including GDPR, customer identity, and sensitive workloads. It also describes Lakebase’s separation of compute and storage for faster ML/AI agent experimentation, alongside Genie’s plain-language data analysis, which can reduce work that otherwise takes months and expensive third parties. Hatton connects these tools with faster, more confident marketing and accurate decision-making in regulated workflows, while also describing Databricks’ inclusive, high-energy culture.


### [What is enterprise intelligence?](https://yomu.fyi/post/what-is-enterprise-intelligence.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 12, 2026

Enterprise intelligence (EI) is presented as an organization-wide capability combining business intelligence, knowledge management, enterprise search and AI to turn structured and unstructured data into decisions and actions. Unlike traditional BI, which centers on dashboards, reports and structured data, EI connects these capabilities through a shared architecture and governed business context. The described stack includes a lakehouse-based data foundation, batch and streaming pipelines, governance, semantics, analytics, search, machine learning, generative AI and a decision layer. Shared definitions such as “active customer” and “monthly revenue” are intended to keep dashboards, queries and AI agents aligned, while maintained context addresses knowledge that becomes stale as the business changes. The result described is a common trusted source from which people, applications and agents can produce insights and initiate actions.


### [Enabling Evolutionary Database Development: Database branching with Lakebase, the conclusion](https://yomu.fyi/post/enabling-evolutionary-database-development-database-branching-with-lak.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Pramod Sadalage, Kevin Hartman
- Published: Jun 12, 2026

The post concludes a series on how copy-on-write database branching in Databricks Lakebase changes team-scale evolutionary database development without changing its underlying methodology. For a team of fifty developers, long-running tier branches and ephemeral feature branches form a parent-linked promotion hierarchy, replacing separately provisioned environment instances and enabling promotion by merge, rollback by repoint, and computable schema divergence. Governance is declared once and inherited per branch, with policies intended to prevent transitions that contradict the parent chain; Unity Catalog captures metadata for attribution and audit. The DBA's role becomes platform engineering, while agents operate inside an executable SCM state machine with documented inputs, outputs, schema validation, and enforced gates. An optional TDD layer adds dedicated roles, acceptance-criterion scenarios, RED-GREEN-REFACTOR cycles, and artifact contracts, and the conclusion presents the resulting workflow as operational for human and agent practitioners.


### [Talk to all your data, wherever it lives](https://yomu.fyi/post/talk-to-all-your-data-wherever-it-lives.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: John Spencer
- Published: Jun 12, 2026

Lakehouse Federation addresses the challenge of reasoning across enterprise data spread among AWS Glue, Snowflake, Oracle, BigQuery, Postgres, and legacy formats without first migrating it. By connecting external sources in place and syncing their metadata into Unity Catalog, Databricks applies shared permissions, lineage, and access controls while preserving source data. Federated comments and descriptions give Genie schema context, while Unity Catalog Semantics lets teams define governed metrics such as ROI once for consistent use across Genie, dashboards, and notebooks. The example connects an AWS Glue marketing database, carries its metadata, defines an ROI metric view, and asks Genie which campaigns led ROI last quarter. The post reports an immediate, accurate answer from live Glue data, and describes planned richer semantics, broader federation, and possible performance gains from managed tables.


### [Unlocking semantics for AI: How Mercedes-Benz Korea built trusted “Talk to Data” at scale](https://yomu.fyi/post/unlocking-semantics-for-ai-how-mercedes-benz-korea-built-trusted-talk.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Sai Yang, Fares Kamal, Alina Kamal, Andreas Jäck, Johannes Laufer, Manuel Culebras
- Published: Jun 11, 2026

Mercedes-Benz Korea piloted a “Talk to Data” architecture that extends its Databricks analytics foundation with a governed semantic layer for enterprise AI, rather than treating the effort as a chatbot project. The design moves Power BI DAX KPI logic into Unity Catalog Business Semantics and Metric Views, keeping sources, joins, measures, dimensions, comments, and synonyms alongside governed Lakehouse data. Genie spaces use curated metric views for domain questions, while Agent Bricks composes persona-based agents, with Unity Catalog enforcing row- and column-level access. An automated DAX-to-Metric-View transpiler parses semantic models, maps tables, generates draft definitions, flags non-automatable measures, and reports conversion gaps. The documented playbook combines gold-layer curation, KPI validation, regression testing, Genie optimization, persona agents, and Databricks Apps; the pilot reports AI answers aligned with established KPI definitions and BI reporting logic.


### [Forward Deployed Engineering: Delivering Business Outcomes with AI](https://yomu.fyi/post/forward-deployed-engineering-delivering-business-outcomes-with-ai.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Jason Martin
- Published: Jun 11, 2026

Databricks is formalizing its Forward Deployed Engineering (FDE) organization to address customers’ shift from migration and data-pipeline requests toward business outcomes with AI. FDE brings Professional Services together around an engineering-led model that embeds engineers with customers, supports modernization and production AI, and works with partners and Databricks R&D. In the cited examples, teams migrated five-plus petabytes of JPMC Consumer and Community Banking Risk data and more than 500 notebooks in four months, while Fox used Lakebase, AI Search, Databricks Apps, and Model Serving to redesign fan experiences. The organization says its engagements use shared OKRs, rapid prototype-to-production delivery, embedded engineering, outcome-aligned commercial options, and global partner coverage, with Fox reporting that Sports AI users spend approximately twice as long in the app.


### [Ingesting the Milky Way: Petabyte-Scale with Zerobus Ingest](https://yomu.fyi/post/ingesting-the-milky-way-petabyte-scale-with-zerobus-ingest.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Aleksandar Tomić, Victoria Bukta, Nikola Obradović, Danilo Najkov, Branko Grbić, Milos Milovanovic
- Published: Jun 11, 2026

The post benchmarks Zerobus Ingest, a managed, serverless, push-based service that writes producer data directly to Delta tables governed by Unity Catalog, against a petabyte-scale telemetry workload. Using NASA’s NEOWISE dataset and Locust, the test modeled fan-in from 2,048 concurrent streams, using Protocol Buffer 2 data over approximately 24 hours. The design replaces static partition-based ordering with stream-connection ordering, allowing heuristic routing across pods, dynamic partitioning, and autoscaling while existing streams drain. It also uses zeroparser, a zero-copy protobuf decoder whose design relies on Rust’s lifetime system and supports dynamic descriptors at about 1 GB/s per CPU core. The test sustained 12 GB/s to one table, ingested 1.04 trillion rows, and reached 1 petabyte within 24 hours; Zerobus Ingest is generally available, with additional APIs on its roadmap.


### [How ERGO Hestia reduced time-to-market with Databricks Lakebase and Model Serving](https://yomu.fyi/post/how-ergo-hestia-reduced-time-to-market-with-databricks-lakebase-and-mo.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Klaudia Ratkowska, Maciej Majewski, Oliver Börner, Alexander Migunov
- Published: Jun 11, 2026

ERGO Hestia redesigned its real-time pricing platform to reduce deployment friction across more than 100 models and 1,000 variables while preparing B2C capabilities. Previously, processed data moved from Databricks through extraction jobs, external Azure PostgreSQL, and a custom caching adapter, creating governance overhead, deployment coordination, and latency spikes during large refreshes. The new architecture uses Lakebase Sync Tables as an online serving layer and Databricks Model Serving Endpoints, keeping data, request logic, and model serving within the lakehouse; Unity Catalog supplies lineage, version tracking, access controls, and audit trails. An incremental migration started with a low-criticality endpoint, measuring 20ms latency and less than 5% CPU utilization at 40 requests per second, before expanding toward larger workloads and the planned decommissioning of PostgreSQL.


### [Welcoming the first cohort of Databricks student fellows](https://yomu.fyi/post/welcoming-the-first-cohort-of-databricks-student-fellows.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Elise Hollowed, Joe Nash, Trang Le
- Published: Jun 11, 2026

Databricks announces its inaugural Student Fellows cohort, selected from more than 5,000 applications submitted by students at hundreds of universities and from dozens of countries. The program targets students who contribute on campus, apply data and AI in practice, and intend to pursue careers in the field, with fellows acting as bridges between academic theory and the real-world scale of the Databricks platform. Five students are profiled, with experience spanning large-scale AI systems, ETL pipelines, computer vision, machine learning platforms, robotics, data modeling, and local retrieval-augmented generation architectures. During the coming academic year, the cohort will receive training from Databricks experts and hands-on experience solving complex data challenges, while building a launchpad for future internship opportunities. The program invites applications for its next cohort in Fall 2026 and points readers to Databricks Free Edition.


### [Geospatial Unbounded: Spatial SQL GA with AI/BI Maps, Delta Sharing, and Iceberg v3](https://yomu.fyi/post/geospatial-unbounded-spatial-sql-ga-with-ai-bi-maps-delta-sharing-and.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Kent Marten
- Published: Jun 11, 2026

Databricks announces Spatial SQL as Generally Available, positioning its platform for geospatial analysis without separate spatial databases, warehouses, and mapping tools. It supports native GEOMETRY columns in Delta or Iceberg, more than 90 OGC-compliant ST\_\* functions, spatial joins, and boolean set operations. AI/BI dashboards can render Geometry and Geography columns as maps, while Genie can generate spatial queries and dashboards and respect Unity Catalog row filters. Geo columns are supported by Delta Sharing, and Databricks can read and write managed Iceberg tables or read externally written Iceberg tables with geospatial types in Iceberg v3. Benchmarks show eight of twelve SpatialBench queries improved since Public Preview, with gains from 20% to 15X, while areal boolean operations are twice as fast on average versus prior versions.


[Newer posts](https://yomu.fyi/company/databricks/page/10.md) · [Older posts](https://yomu.fyi/company/databricks/page/12.md)
