---
title: "Databricks"
description: "Data and AI platform for data engineering, analytics, machine learning, and generative AI."
---

# Databricks
> Data and AI platform for data engineering, analytics, machine learning, and generative AI.

## Articles

### [Beyond dashboards: Introducing Decision Execution Platforms](https://yomu.fyi/post/beyond-dashboards-introducing-decision-execution-platforms.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Marc Solomon, Marcello Pedersen
- Published: Jul 1, 2026

Databricks Forward Deployed Engineering introduces Decision Execution Platforms (DEPs), an enterprise analytics category intended to connect KPI signals, executive decisions, operational execution, and measured outcomes. The proposal addresses workflows in which dashboards reveal problems but meetings, decks, spreadsheets, and messaging threads leave implementation fragmented and impact measurement disconnected. DEPs run the four-stage loop on governed Databricks infrastructure: agents recommend actions, alternatives, predicted impact, and reasoning; approved choices execute through systems of record; and results persist in a Decision Log for continuous learning. Their architecture combines a foundation of Lakebase, Genie, Unity Catalog, Lakehouse, Agent Bricks, and MLflow with an SDK of reusable primitives and a Databricks Apps executive surface. A retailer case used a DEP to unify fulfillment data and enable simulated, controlled rerouting, with scaling aimed at measurable bottom-line and customer-satisfaction outcomes.


### [Forecasting at the speed of modern retail](https://yomu.fyi/post/forecasting-at-the-speed-of-modern-retail.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Ryuta Yoshimatsu, Puneet Jain, Lourdes Angélica Martinez Medina, Lucas Bruand, Dael Williamson
- Published: Jul 1, 2026

Retail and CPG forecasting now spans hundreds of thousands, sometimes far more, time series across fragmented channels, promotions, and short-lived products, making legacy methods and manual exception management difficult. Multi-model forecasting addresses this complexity by evaluating a range of techniques against actual data and selecting the best-performing model for each series, but enterprise-scale experiments require scarce forecasting and distributed-systems expertise. Released in 2024, Databricks’ open-source Many Model Forecasting (MMF) integrates more than 35 statistical, deep-learning, and foundation time-series models and runs on distributed Databricks compute. MMF Agent, built on Genie Code, guides users through data quality, series classification, compute configuration, forecasting, post-processing, and model selection, while Unity Catalog helps it use organizational data context. The workflow is intended to reduce setup from days to hours, improve targeting and accuracy, and make rigorous forecasting more accessible while remaining customizable for technical teams.


### [Celebrating the Winners of the 2026 Built-On Databricks Startup Challenge - Cloned](https://yomu.fyi/post/celebrating-the-winners-of-the-2026-built-on-databricks-startup-challe-cloned.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Andrew Ferguson, Joslyn O'Connell, Sully Clark
- Published: Jul 1, 2026

The 2026 Built-On Databricks Startup Challenge recognized early-stage startups building B2B applications on the Databricks platform, with winners announced at the Data + AI Summit in San Francisco on June 16. VisionHeight won the grand prize for an agentic threat-intelligence platform that maps adversary infrastructure across the Internet while it is being built; Linkup placed second with a production-grade Web Search API delivering results in about two seconds, and Intelo placed third with an agentic workforce for retail merchandising and planning. Clarecast, Gemini Sports, and LakeFusion received honorable mentions for predictive intelligence, football squad planning, and an AI-native data foundation combining MDM, PIM, and LakeGraph. Judges assessed market potential, founding-team caliber, and innovative use of Databricks, while the startup-program offer provides qualifying startups up to $200,000 in credits across Databricks and Neon.


### [From monolith to Lakebase to LTAP: rethinking the database from storage up](https://yomu.fyi/post/from-monolith-to-lakebase-to-ltap-rethinking-the-database-from-storage.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Reynold Xin
- Published: Jun 30, 2026

The post examines why monolithic OLTP databases become fragile and difficult to scale when compute, the write-ahead log (WAL), and data files share one machine. Using Postgres as its primary example, it contrasts local WAL durability and physical cloning with Lakebase, whose stateless Postgres compute externalizes those components into SafeKeeper and PageServer services. SafeKeeper replicates log records across a quorum through Paxos-based network replication, while PageServer applies the WAL and materializes data in cloud object storage. The resulting separation is presented as a way to improve durability, independently scale reads and writes, isolate workloads, and support high availability and branching without physical database clones. LTAP extends the design by making transactional tables directly queryable for analytics from a single governed copy, avoiding CDC or mirroring and allowing transactional and analytical engines to scale independently.


### [How Databricks is turning video into searchable, actionable intelligence](https://yomu.fyi/post/how-databricks-is-turning-video-into-searchable-actionable-intelligenc.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Justin Monaldo, Kacey Hertan, Yvan Aquino
- Published: Jun 26, 2026

Databricks presents video analysis as a data engineering problem for organizations with terabytes of footage that is difficult and expensive to review manually. An app accepts a video and a natural-language prompt, then triggers a Lakeflow job on Serverless GPU Compute to run Meta’s SAM3 segmentation model frame by frame and retain matching moments. Those clips preserve original timestamps and are sent through the Databricks Foundation Model API for summaries that can be written to tables or passed into downstream workflows. In one example, 26 minutes of traffic footage became one minute and 55 seconds of relevant video. The model-agnostic pipeline uses MLflow signatures to support interchangeable or custom models, while event-driven execution and independent serverless GPUs allow concurrent processing without cluster management or idle GPU costs.


### [A Decision Framework for ETL Migration to Databricks](https://yomu.fyi/post/a-decision-framework-for-etl-migration-to-databricks.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Rafael Aielo
- Published: Jun 26, 2026

An ETL migration with hundreds of stored procedures, schedulers, scattered permissions, and a warehouse renewal deadline needs workload-by-workload decisions rather than a single rewrite strategy. The framework assigns work among Databricks SQL, Spark Declarative Pipelines (SDP), and PySpark or Spark SQL notebooks. SQL tasks suit single statements, Unity Catalog-governed stored procedures handle procedural logic, SDP manages dependencies, retries, quality constraints, and batch-plus-streaming, while notebooks cover complex logic, ML feature engineering, integrations, and large or tightly controlled Spark workloads. It recommends four phases—assessment, quick wins, modernization, and optimization—using profiling, side-by-side validation, and parallel runs before retiring legacy systems. Migration tools can automate 60–80% of initial conversion, but architecture choices remain essential: the goal is consolidating orchestration, metadata, lineage, permissions, and validation rather than reproducing technical debt.


### [How the English Office for Students leverages Databricks to enhance higher education standards and drive better student outcomes](https://yomu.fyi/post/how-the-english-office-for-students-leverages-databricks-to-enhance-hi.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Kacey Hertan
- Published: Jun 26, 2026

The Office for Students, which regulates more than 400 higher education providers in England, modernised its data and analytics environment after a legacy platform could no longer handle growing volumes, varied sources, or emerging analytical demands. Its data spans millions of student records collected over 15 to 20 years, and a workflow processing about 300 million records took eight hours. Moving to Databricks consolidated structured, qualitative, and near-live data with analytics and AI workflows, while Unity Catalog added lineage, access controls, and security patterns for governed use. Genie Code reduced a student segmentation analysis from at least two weeks for two analysts to half a day, and a provider-registration triage proof of concept flags missing submissions earlier. The organisation frames AI as decision support rather than decision-making, keeping humans responsible for regulatory judgments.


### [From test bench to lakehouse: how AVL modernizes measurement data analytics with Impulse](https://yomu.fyi/post/from-test-bench-to-lakehouse-how-avl-modernizes-measurement-data-analy.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Dr. Thomas Bonfert, Jonathan Bräuer, Fabian Ade, Maxim Hammer, Florian Gorzitzke, David Crescence, Christa Simon, Jörg Zimmermann, Hannes Schneider
- Published: Jun 25, 2026

AVL’s Lakehouse for Measurement Data addresses the scale, reproducibility, and governance limitations of desktop tools and isolated scripts used for automotive measurement analysis. Built on Databricks, the platform ingests ASAM MDF4 and other files into a Medallion Architecture, applies configurable DQX quality rules in a hierarchical Silver model, and uses Impulse to compile declarative Python TSAL expressions into distributed Spark execution. Engineers can select channels, create virtual signals with alias resolution, unit conversion, time alignment, and interpolation, define events, and compute duration- or distance-weighted aggregations in about 10 lines of Python. Impulse supports Gold-layer reporting, ad-hoc Spark DataFrames, and ML feature matrices, with Unity Catalog governance and Workflow orchestration. AVL reports reducing analysis time from days to minutes, processing many recordings per run, lowering infrastructure costs versus on-premises solutions, and enabling self-service, reproducible, standardized analysis.


### [What To Look For in a Serverless Database for AI Applications](https://yomu.fyi/post/what-to-look-for-in-a-serverless-database-for-ai-applications.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 25, 2026

Serverless databases are presented as a baseline for AI applications whose traffic can be volatile, idle for long periods, or spike when agents fan out queries. The guide distinguishes managed serverless systems from autoscaling products by focusing on architectural separation of compute and storage, true scale-to-zero, cold-start behavior, connection handling, pricing, performance, portability, governance, and AI capabilities such as vector search. It recommends evaluating both low- and high-utilization costs, published warm-up times, tail latency (p95/p99), and built-in pooling or HTTP/Data APIs for high-concurrency agents and serverless functions. The text positions Lakebase as an example that combines serverless Postgres, shared lakehouse storage, and Unity Catalog governance, and cites reported cost and management reductions from a 2025 study while noting that provisioned deployments may suit continuously high-throughput workloads.


### [What Is Serverless PostgreSQL?](https://yomu.fyi/post/what-is-serverless-postgresql.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 25, 2026

Serverless PostgreSQL is presented as a fully managed cloud database model that decouples compute and storage, allowing each to scale independently with demand. Traditional deployments require teams to size infrastructure, manually manage scaling, and absorb costs from idle capacity. In serverless systems, the provider provisions compute on demand, can suspend it when idle, and bills according to active usage; scale-to-zero may introduce cold-start latency. The architecture can also support database branching through copy-on-write, creating isolated environments without duplicating data. The article distinguishes this model from lakebase architecture, which combines transactional and analytical workloads on a shared foundation using decoupled compute, durable object storage, log-based storage systems, and orchestration.


### [The Rise of Sports Intelligence: How the Lakehouse Turns Tracking Data into Competitive Advantage](https://yomu.fyi/post/the-rise-of-sports-intelligence-how-the-lakehouse-turns-tracking-data.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Corey Abshire, Kush Patel, Nick Ragonese
- Published: Jun 24, 2026

Professional basketball’s Hawk-Eye SkeleTRACK feed produces roughly 22,620 positional updates per second—about 65 million records per 48-minute game—yet teams often cannot turn that volume into timely, trusted decisions. The integration gap comes from separate vendors for tracking, wearables, video, scouting, and medical data, alongside calibration differences, weak provenance, and compute limits. The Databricks Data + AI Platform is presented as a governed lakehouse that ingests feeds with Lakeflow, refines them through medallion layers, and uses Unity Catalog for lineage, access control, and auditing. Models for shot probability, injury risk, and fatigue can run alongside serving and custom applications, with Lakebase supporting sub-second interactive queries. Applications include proactive load management, real-time coaching intelligence, and enriched broadcast or fan experiences across tracking-rich sports.


### [How Daikin Applied Americas builds consistent data pipelines at scale with Genie Code](https://yomu.fyi/post/how-daikin-applied-americas-builds-consistent-data-pipelines-at-scale.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Trent Lezer, James VanGordon
- Published: Jun 24, 2026

Daikin Applied Americas needed to scale reliable data pipelines across growing analytics and AI use cases involving operational, manufacturing, and service data while coordinating development across teams. It adopted Databricks Genie Code within a structured operating model, using Unity Catalog context, reusable MECE skills, and explicit checkpoints across Bronze, Silver, and Gold layers to guide planning and execution. The framework defines competencies such as source grain, transformation patterns, canonical alignment, governance, and business-entity modeling, moving standards out of long prompts and into the development environment. The team reports that pipelines that once took days to prototype could be generated in minutes, with faster iteration, more consistent outputs, less structural correction, reduced architectural drift, and greater trust in AI-assisted results.


### [What are Dashboards?](https://yomu.fyi/post/what-are-dashboards.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 24, 2026

A dashboard is a visual interface that combines metrics, KPIs, and visualizations from one or more sources so users can judge performance against a goal at a glance. It acts as an organized, live visual layer over databases, warehouses, cloud applications, spreadsheets, or other sources, rather than being merely a collection of individual charts or reports. A typical flow sends source data through a query or pipeline into a visualization layer; modern dashboards may query underlying data directly and refresh on a schedule or in real time, avoiding stale copied data and definition drift associated with older extract-based architectures. Filters and drill-downs support investigation, while dashboard types—operational, analytical, strategic, and tactical—differ by audience, time horizon, and update frequency; usefulness ultimately depends on a clear purpose, defined audience, trusted data, and enough context for action.


### [What if the answer was already in your data?](https://yomu.fyi/post/what-if-the-answer-was-already-in-your-data.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Bryan Smith, Morgan Wilkie, Kaitlin Ryan
- Published: Jun 24, 2026

Kythera Labs is building an AI-native healthcare strategy platform on Databricks to give health systems access to strategic intelligence that historically required specialized analysts or consulting firms. Its foundation converts 339 billion medical and prescription claims covering more than 300 million patients into governed, event-based data, resolving providers, harmonizing codes across 130 vocabularies, and reconstructing patient journeys. Healthcare Strategy Agent, built with Agent Bricks, lets executives ask questions such as where oncology referrals are going and receive analyses of leakage, competing providers, physicians, and reimbursement opportunity in minutes. A Louisiana health system went live within ten days and reported 150% greater visibility into encounters, 12% more keepage, 22% less leakage, and $3.8 million in estimated annualized retained-encounter value. Unity Catalog, Lakebase, Delta Lake, Delta Sharing, and serverless infrastructure provide shared governance, lineage, access controls, and operational integration.


### [Databricks positioned highest in execution and furthest in vision for the second consecutive year in Gartner Magic Quadrant](https://yomu.fyi/post/databricks-positioned-highest-in-execution-and-furthest-in-vision-for.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Craig Wiley, Kasey Uhlenhuth, Kayli Berlin, Cynthya Peranandam
- Published: Jun 24, 2026

Databricks says Gartner positioned it highest for execution and furthest for vision in the Magic Quadrant for the second consecutive year. The post connects this recognition to a category reclassified from “Data Science and Machine Learning” to “AI Platforms for Data Science and Machine Learning,” and argues that agentic applications require enterprise data, governance, observability, and business context. Databricks presents a unified approach combining the lakehouse, Lakebase, Agent Bricks, Unity Catalog, and Unity AI Gateway to build, monitor, and govern agents, models, data, apps, and tools. Reported examples include YipitData’s 20x increase in company coverage with 92–95% tagging accuracy, Block’s unified AI and data estate, and Novo Nordisk’s attribution of more than $157 million in net new value to governed clinical-trial optimization.


### [Genesis Workbench: A blueprint for industry AI in life sciences, powered by Databricks and NVIDIA](https://yomu.fyi/post/genesis-workbench-a-blueprint-for-industry-ai-in-life-sciences-powered.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Mark Lee, Srijit Nair
- Published: Jun 23, 2026

Genesis Workbench is an open blueprint for a Databricks-native life-sciences application that brings computational drug-discovery stages into one governed workbench. It combines Unity Catalog governance, MLflow tracking, Model Serving, serverless GPU compute, Databricks AI Search, and NVIDIA technologies including CUDA-X libraries, Parabricks, BioNeMo tools, GenMol, and Proteina-Complexa. Independent modules cover genomics, single-cell analysis, large- and small-molecule workflows, and model fine-tuning, with handoffs spanning gene-to-sequence resolution, structure prediction, docking, ADMET, and candidate ranking. A point-and-click React interface supports bench scientists, while declarative workflow generation and MCP exposure let pipelines and external clients use the workbench; inference runs on GPU endpoints inside the governed workspace without runtime external API dependencies. The stated aim is to let teams move from disease hypotheses to ranked therapeutic candidates on their own data, with a roadmap for automated workflow generation, BioNeMo Skills integration, and additional MCP services.


### [Guide to Agentic Systems and AI Agents](https://yomu.fyi/post/guide-to-agentic-systems-and-ai-agents.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 23, 2026

Agentic AI systems are goal-directed software platforms that perceive context, plan and execute multi-step workflows, and adapt based on outcomes with minimal human intervention. The guide distinguishes them from traditional and generative AI, defining agents, broader system architecture, and the role of LLMs as reasoning cores connected to memory, APIs, databases, and other tools. It describes a perceive-reason-act-learn loop, multi-step planning, external tool integration through interfaces such as the Model Context Protocol (MCP), and orchestration patterns for coordinating specialized agents. Production concerns include retries, queues, observability, permissions, privacy, logging, and human escalation, while stated risks include reward-hacking, unintended actions, and explainability gaps. It identifies repetitive, data-rich workflows with clear success criteria and bounded error consequences as the best current enterprise candidates.


### [Top 10 AI Business Solutions Driving Company Growth](https://yomu.fyi/post/top-10-ai-business-solutions-driving-company-growth.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 23, 2026

The article identifies ten AI business solution categories presented as growth drivers, while arguing that value is concentrated in workflows where AI changes the economics of work. It frames successful adoption around three conditions: clean, governed data; process-first use-case selection; and governance designed in from the start. Examples include customer-service agents, forecasting, personalization, intelligent process automation, and supply-chain optimization; the text says customer service accounts for 40% of top use cases, while data quality accounts for roughly 75% of what makes an AI solution work. It also describes productivity, automation, and business reimagination as distinct value paths, including a payments-data forecasting product that became an eight- to nine-figure annual revenue stream. The conclusion favors unified platforms that connect data, analytics, AI, governance, and agentic workflows.


### [End-to-End RAG Workflow: How Retrieval Augmented Generation Works](https://yomu.fyi/post/end-to-end-rag-workflow-how-retrieval-augmented-generation-works.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 23, 2026

Retrieval Augmented Generation (RAG) connects a large language model to external knowledge at inference time, addressing outdated training data and access to proprietary or real-time information. The workflow ingests and normalizes sources, splits documents into chunks, embeds them in a vector store, retrieves context, assembles a prompt, and generates an answer. Semantic search can be combined with BM25 keyword search through reciprocal rank fusion, while reranking can improve precision; the same embedding model must be used during ingestion and querying. The guide presents evaluation and deployment considerations, including separate measurement of retrieval precision and generation faithfulness, versioning, monitoring, and containerized components. It identifies poor retrieval as the most common failure mode and explains that RAG reduces, but does not eliminate, hallucinations.


### [What is Vector Search?](https://yomu.fyi/post/what-is-vector-search.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 23, 2026

Vector search retrieves results by comparing embeddings that represent meaning across text, images, audio, and other content rather than matching exact words. A model creates embeddings, an index stores them for fast similarity search, and a query embedding is matched against the index using nearest-neighbor methods. Exhaustive k-nearest neighbor search can become too slow at millions of items, so production systems commonly use approximate nearest neighbor search, trading some precision for speed. The guide positions vector search behind semantic search, RAG, recommendations, and multimodal or cross-language retrieval, while hybrid search combines dense and sparse vectors, keyword results, metadata filtering, and reranking to improve reliability. Quality depends on embeddings, filters, index freshness, and infrastructure, with vector search requiring more memory and compute; Databricks AI Search is presented as a managed service supporting these capabilities and Unity Catalog governance.


[Newer posts](https://yomu.fyi/company/databricks/page/7.md) · [Older posts](https://yomu.fyi/company/databricks/page/9.md)
