---
title: "Databricks"
description: "161 posts about Databricks, summarised, each linking to the original."
---

# Databricks
> 161 posts about Databricks, summarised, each linking to the original.

## Articles

### [The audience is the asset. Media finance teams need to understand them to protect the margin.](https://yomu.fyi/post/the-audience-is-the-asset-media-finance-teams-need-to-understand-them.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Elena Tesser
- Published: Jul 28, 2026

Media finance teams are tasked with understanding how audience value flows across subscriptions, advertising, and content investment so the business can protect margin. Streaming has turned one wholesale audience into multiple monetization strategies, while ad-supported tiers now account for 59% of new streaming sign-ups, making timely measurement and pricing analysis more important. Ontology preserves the meaning and context of figures as audiences, titles, channels, and business conditions change, distinguishing a correct answer from one that is merely accurate. Databricks Genie is described as a governed, data-smart AI coworker that answers sourced natural-language questions, learns from interactions, and shows its work; Genie-powered apps let hundreds of DIRECTV analysts and leaders query more than 1,200 customer-level attributes, while people retain decision authority.


### [From prototype to production: High QPS for Databricks AI Search](https://yomu.fyi/post/from-prototype-to-production-high-qps-for-databricks-ai-search.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Adam Gurary, Sheng Zhan, Ankit Vij, Vadim Antonov, Yu-Ju Huang, Dima Kotlyarov
- Published: Jul 28, 2026

Databricks AI Search now offers generally available high-QPS scaling for Standard endpoints, addressing production workloads such as search bars, recommendations, and real-time entity resolution. Operators set a human-readable target\_qps value when creating or updating an endpoint through the Python SDK, REST API, or UI, while Databricks provisions the required compute without manual replica counts, node sizing, or load balancers. Existing Unity Catalog governance and Delta Sync remain in place, and scaling progress is exposed through scaling\_info as it moves from SCALING\_CHANGE\_IN\_PROGRESS to SCALING\_CHANGE\_APPLIED; endpoint observability shows requests per second, latency, and health. Service principal authentication is recommended for high-QPS traffic, while personal access tokens are capped at a few tens of QPS, and new capacity applies when an index is created or synced.


### [Get Started with Genie One: Top AI Cowork Use Cases for Business Users](https://yomu.fyi/post/get-started-with-genie-one-top-ai-cowork-use-cases-for-business-users.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Cynthya Peranandam, Sydney Sundell
- Published: Jul 28, 2026

Genie One is presented as a data-smart, agentic coworker for business users who need work completed across calendars, CRMs, data warehouses, ticketing systems, and shared documents rather than through a standalone chatbot. The guide explains four workflows: recurring business reviews, meeting preparation and follow-up, knowledge-work document automation, and operational monitoring with alerts. Users connect relevant sources, provide templates or natural-language rules, create reusable skills, test outputs, schedule routines, and refine instructions over time; the monitoring workflow can compare metrics with defined thresholds and email alerts. The post recommends starting with one narrow recurring task, connecting its required sources securely with data or IT partners, and iterating toward a library of saved skills that can be used across the organization.


### [The EU Digital Product Passport: a traceability deadline](https://yomu.fyi/post/the-eu-digital-product-passport-a-traceability-deadline.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Daniel Dahlin
- Published: Jul 27, 2026

The EU’s Ecodesign for Sustainable Products Regulation (ESPR) will require in-scope products sold in the single market to carry a machine-readable Digital Product Passport (DPP), with batteries first from February 2027. The passport links a unique identifier and data carrier to lifecycle information that must remain accurate, while operators retain responsibility for data stored on their own backend, tiered access, supplier inputs, and any Registry registration. Implementation is primarily a data-integration and governance problem, spanning tier-N suppliers, per-unit operational records, lineage, analytics, AI, and controlled sharing rather than QR-code generation. Databricks components including Lakebase, Unity Catalog, Lakeflow Spark Declarative Pipelines, Databricks Apps, and OpenSharing are presented with a battery-focused Solution Accelerator as an operator-side reference architecture. The stated benefits extend beyond market access to faster recalls, sourcing-risk visibility, and lineage-based sustainability reporting, but technology alone does not establish compliance.


### [How the FDA Built an AI Platform That 85% of Its Staff Now Use Daily](https://yomu.fyi/post/how-the-fda-built-an-ai-platform-that-85-of-its-staff-now-use-daily.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Molly Just-Behr
- Published: Jul 23, 2026

The FDA built ELSA, a generative AI platform for its 16,000 staff, on Halo, a governed Databricks data foundation created to address fragmented systems across eight centers. Those centers had separate chatbots and data stores; consolidating 50 to 60 sources enabled faster sharing, real-time streaming, and centralized access controls through Unity Catalog. Within roughly two months, ELSA adoption rose from less than 1% to 85%, while staff began building hundreds of agents weekly from standard operating procedures, regulatory guidance, and center-specific documents. MCP servers layered over Unity Catalog make governed data and tooling accessible beyond data scientists, and Databricks ML and NLP capabilities through MLflow extracted starting materials and product-supplier-manufacturer relationships from millions of submission pages. A reviewer can now request grounded starting-material information for a drug application in about three minutes instead of days, while the FDA adapts center-specific MCP tools and extends the model across its organization.


### [Provisioning for the Agentic Era: How Databricks Built a Self-Serve Infrastructure Vending Machine](https://yomu.fyi/post/provisioning-for-the-agentic-era-how-databricks-built-a-self-serve-inf.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Evan Pandya, Greg Wood, Joel Thomas
- Published: Jul 23, 2026

Databricks built the Field Engineering Vending Machine (FEVM) to replace shared environments with isolated, centrally governed infrastructure for its growing field organization. Users or agents describe a use case and configuration details, receiving an environment assembled through a React frontend, Python backend, Databricks Apps, Terraform, Git Runner, and Lakebase across AWS, Azure, and GCP. FEVM supports templates and add-ons such as Lakebase, notebooks, and pre-packaged assets from UC Volumes, applies configurable time-to-live policies, and sends Slack notifications for provisioning, expiration, and deletion. It also manages shared resources independently, including catalogs that persist across workspace lifecycles, and provides administrative controls for limits and audits. At the time described, it managed more than 2,600 active deployments across three clouds and over 5,000 active users, after handling nearly 1,200 requests during one internal event.


### [Why A Frontier Data Agent Outperforms General Coding Agents in Quality and Cost](https://yomu.fyi/post/why-a-frontier-data-agent-outperforms-general-coding-agents-in-quality.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: The Databricks AI Research Team
- Published: Jul 23, 2026

An evaluation compared Genie Code with three widely used coding agents from major AI labs on 401 real internal data tasks, using a shared 20-minute wall-clock budget and independent grading. Genie Code achieved 76.6% accuracy at an estimated $0.55 per task, outperforming rivals at 55.9–72.1% accuracy and $0.91–$1.16 per task. Its advantage is attributed to semantic search across catalog and workspace assets, persistent memory of tables and business logic, and enterprise-context understanding, which reduce exploratory tool use. On discovery-heavy tasks, general-purpose agents often wandered through large workspaces or timed out on inefficient scans, while Genie Code averaged 8.3 tool calls per task. The benchmark found no accuracy-cost trade-off, although the authors note that Genie Ontology was disabled and expect it to strengthen the advantage.


### [Connect Amazon S3 data to Databricks with Delegated IAM Permissions](https://yomu.fyi/post/connect-amazon-s3-data-to-databricks-with-delegated-iam-permissions.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Gordon Wang
- Published: Jul 23, 2026

Databricks introduces a simpler flow for connecting Amazon S3 to Unity Catalog through an external location, addressing the manual work previously required to establish governed read and write access. Instead of writing 140-line IAM trust policies, configuring bucket permissions, deploying CloudFormation templates, and switching between consoles, users specify a bucket and access level, then verify permissions in AWS. AWS IAM temporary delegation lets Databricks provision a least-privilege IAM role, configure its cross-account trust policy, register the external location, and enable Auto Loader and File Events. The delegated authorization is time-bounded and expires after setup; provisioning actions are logged in AWS CloudTrail, and users without sufficient permissions can request access from an AWS administrator in the flow. The result is a single-session setup intended to reduce configuration errors while supporting governed S3 access for ingestion, pipelines, analytics, and LTAP workloads.


### [Introducing AI spend controls with Unity AI Gateway](https://yomu.fyi/post/introducing-ai-spend-controls-with-unity-ai-gateway.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Kevin Stumpf
- Published: Jul 23, 2026

Unity AI Gateway now offers AI Spend Controls, extending existing cost visibility with proactive budget alerts across models and workloads. The feature supports budgets at user, use-case, workspace, and account levels, plus shared and per-user thresholds that can trigger email alerts or enforce hard caps by stopping requests after a limit is exceeded. Configuration starts in account settings under Usage and Budgets, where administrators select Unity AI Gateway, optionally scope workspaces and resource tags, and define monthly limits and recipients. Budget status and trends are available in the Cost section, while customizable Cost Analytics dashboards use Unity Catalog system tables to attribute DBU and model-provider costs by identity, workspace, endpoint, tags, model, provider, and request tags. The release positions Databricks budgets, Unity AI Gateway, and Unity Catalog as a combined governance layer for controlling AI access, usage, and spend.


### [Simplify AI agent orchestration with Lakebase Postgres](https://yomu.fyi/post/simplify-ai-agent-orchestration-with-lakebase-postgres.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Li Yu, Michelle JanneyCoyle, Jon Cormack, Yarri Bryn, Alec Sorensen, Darshana Nair
- Published: Jul 22, 2026

CLA and Databricks built a production document-processing application for auditing that reduces extraction time from hours to minutes without compromising quality, using Databricks-native services including Lakebase Postgres, Databricks Apps, Lakeflow Jobs, MLflow, and Unity Catalog Volumes. Lakebase serves as the orchestration layer’s single source of truth for tasks and execution attempts, coordinating long-running work, retries, leases, priorities, rate limits, costs, and status visibility. The queue uses Postgres patterns including FOR UPDATE SKIP LOCKED, priority and FIFO ordering, expiring leases for crash recovery, and database-backed concurrency controls. Databricks Jobs process PDFs through intelligent document processing and vision/LLM calls, while MLflow Tracing records execution and cost details and dashboard updates combine fast Postgres data with slower billing queries. In production, this architecture avoids external brokers and schedulers while providing durable task management, real-time visibility, and per-task cost attribution.


### [The last mile: why great first-party data still doesn't make great marketing](https://yomu.fyi/post/the-last-mile-why-great-first-party-data-still-doesn-t-make-great-mark.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Michael Burton, Katy Yuan
- Published: Jul 21, 2026

Modern data platforms and sophisticated marketing systems can still fail to produce timely customer experiences when no bridge connects first-party data to campaign activation. The post explains Scott Brinker’s “composable canvas,” a five-ring architecture centered on a unified data core, with semantic layer, CaaS, decisioning, and apps and agents operating on shared data without repeated movement. In contrast, batch files, dashboards, and disconnected teams can leave autonomous agents unable to trigger campaigns or propensity models unused. It recommends closing this last mile through activation-ready data architecture, self-service marketing analytics, and one narrowly scoped AI agent tied to a measurable campaign outcome. Examples in the post include a 4X model conversion rate, a four-week account scoring launch, and grocery offer flows that changed weekly preparation from hours to minutes.


### [How Dow Built a Carbon Footprint Ledger on Databricks to Accelerate Sustainability at Scale](https://yomu.fyi/post/how-dow-built-a-carbon-footprint-ledger-on-databricks-to-accelerate-su.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Jesse Grekowicz, Tim Licquia, Varun Mahajan
- Published: Jul 21, 2026

Dow built the Carbon Footprint Ledger (CFL) on Databricks to provide a transparent, verifiable way to track and communicate product carbon footprints amid customer Scope 3 commitments and tighter environmental disclosure requirements. The system combines carbon-accounting methodology assured against ISO 14067 and the GHG Protocol Product Standard with a calculation engine and certificate-issuing ledger that maintains residual balances. Within Dow’s Integrated Data Hub, Apache Spark unifies supply-chain, sustainability, operations, and other source data, while Unity Catalog governs access and Delta tables provide ACID transactions, time travel, and efficient upserts. An advanced optimization model, developed and deployed with MLflow, identifies lowest-greenhouse-gas production pathways and sends versioned, monitored results through the pipeline with lineage tracking. The implementation reduced PCF processing from weeks to a fraction of that, enabled portfolio-scale calculations, and supports verifiable certificates and commercial use of decarbonization investments.


### [Why R&D Data Belongs in the Lakehouse - and Why Agents Need It There](https://yomu.fyi/post/why-r-d-data-belongs-in-the-lakehouse-and-why-agents-need-it-there.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Sebastian Eberhardt, Dominik Bentele, Jonathan Bräuer
- Published: Jul 21, 2026

cellcentric describes how it built a governed lakehouse foundation for R&D data so industrial AI agents can reason over engineering context, not isolated source-system records. On Azure and Databricks, the Data Hub combines Unity Catalog governance, Lakehouse Federation, Delta Sharing, Declarative Automation Bundles, and an MCP interface for agents. Its Fuel Cell Passport unifies five enterprise systems, models seven hierarchy levels, refreshes daily, and uses a state-based temporal model for point-in-time configuration and rework-history questions. The platform treats context coverage as a data-quality metric, with 27 published products averaging 90% column-comment coverage, while authenticated user identity and Unity Catalog authorization constrain agent access. For many R&D and process-development workflows, work that once took weeks now ships in days, though complex multi-source investigations still require substantial domain review.


### [Announcing the Public Preview of Discover and Domains, powered by Unity Catalog](https://yomu.fyi/post/announcing-the-public-preview-of-discover-and-domains-powered-by-unity.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Stef Bran Melendez, Kelly Albano
- Published: Jul 21, 2026

Databricks announces the Public Preview of Domains and the Discover page, powered by Unity Catalog, to help people and agents identify relevant, high-quality, and safe data and AI assets. Domains organize tables, dashboards, notebooks, queries, metric views, Genie Agents, and apps by business structure, while Discover provides an internal marketplace with search, certification signals, popularity and trending indicators, and AI-powered recommendations. Data stewards can create domains and subdomains, certify assets, add descriptions and contacts, and curate page sections and pinned content. Domains extend Unity Catalog Semantics and feed Genie Ontology, giving agents business context for narrowing retrieval, prioritizing trusted assets, and interpreting metrics within each user’s existing permissions. The features are available in Public Preview through a Databricks workspace, where organizations can create domains and curate assets.


### [Branching databases like code: a CI/CD pattern for Lakebase, in production at Glaspoort](https://yomu.fyi/post/branching-databases-like-code-a-ci-cd-pattern-for-lakebase-in-producti.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Hadi Farhat, Gideon Spierings, Ricardo de Vries, Raymon Veldman
- Published: Jul 20, 2026

Glaspoort replaced a sprawl of outdated, one-off BI reports with a custom application that combines Databricks analytical tools with Lakebase, a serverless Postgres OLTP database. The application keeps curated lakehouse data in a read-only schema while application state occupies a separate schema, and supports development, acceptance, and production environments. Its CI/CD pattern treats database branches like code: long-lived environments branch directly from production, each pull request receives an ephemeral production-shaped branch, migrations are replayed, and a staging app image runs the full test suite against it. This topology makes periodic resets cheap, avoids deleting dependent branches, and preserves fresh test data; the team chose stacked promotion for velocity, retaining a separate crisis pipeline for urgent fixes.


### [The three ways AI unlocks transformation in Retail, Travel, and Consumer Goods](https://yomu.fyi/post/the-three-ways-ai-unlocks-transformation-in-retail-travel-and-consumer.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Rob Saker
- Published: Jul 20, 2026

The piece argues that retail, travel, and consumer-goods companies face one problem in three forms: signals are not trusted, arrive too late, or cost too much to process at scale. It contrasts business intelligence, which depends on structured schemas, predefined questions, and dashboards, with AI systems that read unstructured data, reason probabilistically across signals, and connect decisions to action. Examples include reported improvements from Harmons’ shelf scanning, faster consumer-insight work at a health and hygiene company, and travel applications spanning maintenance, pricing, and concierge services. Its proposed architecture combines broad ingestion and open storage with governance, evaluation, model and agent controls, and applications that operate on a substrate, presenting current, coherent data as the foundation for organizations that can act.


### [Tech builds on AI. Finance protects the margin.](https://yomu.fyi/post/tech-builds-on-ai-finance-protects-the-margin.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Madelyn Mullen
- Published: Jul 17, 2026

AI-native tech companies must protect unit economics as agents accelerate changes in compute consumption, pricing, and revenue recognition, while gross margins remain below classic software levels. Finance teams built around extracts, spreadsheets, and monthly reconciliation can miss repricing changes, metering errors, and compute-commitment risk. The proposed foundation is an evolving ontology that keeps product, plan, usage, and cost meanings current, with Stripe data entering Unity Catalog through OpenSharing and Lakebase providing transactional Postgres on the lakehouse. Genie One uses that ontology to answer governed, sourced questions about gross margin, consumption revenue at risk, and compute spend, while people retain decision authority. The post describes organizations using Databricks to consolidate reporting, forecasting, workflows, and finance applications, positioning a shared data-and-AI platform as the path from an initial answer to an ongoing finance platform.


### [Your AI is ready. Your data foundation probably isn’t](https://yomu.fyi/post/your-ai-is-ready-your-data-foundation-probably-isn-t.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: CIO.com
- Published: Jul 16, 2026

Cushman & Wakefield’s enterprise AI program addresses fragmented experiments, disconnected data, and uneven organizational maturity across a 53,000-person workforce. Over four years, Chief Digital and Information Officer Sal Companieh used a product operating model, business-linked accountability, co-created investment decisions, and shared architecture standards to build a common foundation while preserving business-unit flexibility. Databricks supports that strategy as a partner and platform, with its intelligence layer and Genie helping employees query trusted data in natural language, examine quality and governance, and monitor compliance. The company says the time from idea to outcome has fallen from months to days, while client and acquisition onboarding has materially accelerated. Companieh identifies human behavior, education, and trust—not technology alone—as essential to making change durable.


### [From experiment to insight: how Dotmatics Luma and Databricks make AI-ready science a reality](https://yomu.fyi/post/from-experiment-to-insight-how-dotmatics-luma-and-databricks-make-ai-r.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Ryan Bernhardt, Michael Fritz
- Published: Jul 16, 2026

Scientific workflows generate data across instruments, teams, and partner networks, but siloed systems can strip away metadata, lineage, and experimental context. Luma, Dotmatics’ scientific intelligence platform, captures outputs continuously and harmonizes them into structured, FAIR-compliant records, while Databricks supplies scalable storage, governance, and enterprise data and AI infrastructure. Running natively on Databricks, Luma preserves a continuous digital thread across experiment design, acquisition, analysis, reporting, and sharing, with Delta Sharing supporting governed exchange with collaborators. Chromatography illustrates the approach: connectivity across vendor systems and Virscidian’s Analytical Studio automate LC/MS processing while retaining context and adding dashboards, registration, and management tools. In one pharmaceutical deployment, Luma connected about 1,500 of more than 5,000 instruments across four LC/MS vendors, enabling cross-vendor performance trending, unified purity analysis, and a foundation for AI and machine learning.


### [The skills gap behind agentic AI — and how Databricks is closing it with a new context engineer certification and agent trainings](https://yomu.fyi/post/the-skills-gap-behind-agentic-ai-and-how-databricks-is-closing-it-with.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Rachel Canetta, Trang Le
- Published: Jul 16, 2026

Databricks introduced the Databricks Certified Context Engineer Associate beta exam to validate skills for building reliable, production-grade AI agent systems as organizations scale agentic AI. Context engineering is presented as the practice of curating, maintaining, and filtering tokens, memory banks, and tool parameters so an LLM can solve a task; the certification is intended to benchmark that technical skill set. Databricks also added AI Agent Fundamentals, Building Retrieval Agents on Databricks, and Agent Evaluation, covering agent reasoning, retrieval-augmented architectures, and systematic performance testing and improvement. Its AI-first certification-prep guide is available on every certification page, works with free tiers of major LLMs, includes guardrails and disclaimers about hallucinations, and shows candidates how to use Free Edition for hands-on practice; registration for the Context Engineer Associate is open, with the first exam scheduled for July 29, 2026.


[Newer posts](https://yomu.fyi/topic/databricks.md) · [Older posts](https://yomu.fyi/topic/databricks/page/3.md)
