---
title: "Databricks"
description: "Data and AI platform for data engineering, analytics, machine learning, and generative AI."
---

# Databricks
> Data and AI platform for data engineering, analytics, machine learning, and generative AI.

## Articles

### [Addressing HR's widening capacity gap with AI](https://yomu.fyi/post/addressing-hr-s-widening-capacity-gap-with-ai.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Paurul Yadav, Soumya Dash, Bryan Smith
- Published: May 8, 2026

HR teams face a widening capacity gap as strategic expectations, complex employee issues, workforce volatility, skills shortages, and demands for personalized support collide with largely unchanged headcount and tools. The article presents AI transformation as an incremental journey: first establish a secure Employee 360 from structured and unstructured enterprise data, then build reusable workforce insights, augment workflows with human oversight, and progress toward broader transformation. It emphasizes data governance, including access controls, auditing, quality, standardization, and reliable interpretations, while noting that trust has limited AI’s business impact so far. MathCo and Databricks support this roadmap through NucliOS, whose Data Studio, AI Studio, and Decision Studio environments connect governed data, explainable models, feedback loops, and decision applications; Databricks supplies the lakehouse foundation, lineage, quality checks, and privacy-compliant access.


### [MCP Marketplace brings real-time intelligence to agentic applications](https://yomu.fyi/post/mcp-marketplace-brings-real-time-intelligence-to-agentic-applications.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Roman Ostrovski, Harish Gaur, Antoine Amend
- Published: May 8, 2026

The MCP Marketplace connects agentic applications with real-time external intelligence alongside enterprise data. The problem appears in use cases such as loan approval, where historical internal records omit market conditions, updated credit signals, property changes, and competitor activity, making manual research a bottleneck. Databricks Marketplace provides governed access to MCP servers from You.com, Moody’s, and Cotality, while Unity Catalog authenticates connections and tracks access and lineage; Lakebase stores state, decisions, and audit trails across multi-step workflows. Examples show agents combining internal data with web research, credit ratings and sector outlooks, or property-resolution and mortgage signals before surfacing decisions for human review, including a commercial-loan flow with recorded sources, timestamps, and approver.


### [Pushing the Frontier for Data Agents with Genie](https://yomu.fyi/post/pushing-the-frontier-for-data-agents-with-genie.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: The Databricks AI Research Team
- Published: May 8, 2026

Genie is Databricks’ data agent for complex questions across structured enterprise assets—tables, dashboards, and notebooks—and unstructured sources including workspace files, Google Drive, and Sharepoint. Unlike coding agents operating in static environments, it must discover relevant assets at enterprise scale, determine authoritative knowledge from potentially contradictory sources, and handle questions without verifiable tests or guaranteed answers. It combines specialized knowledge search using semantic context and metadata, parallel thinking across sampled trajectories, and Multi-LLM orchestration with optimized prompts for distinct sub-agents. On an internal benchmark of real-world data-analysis tasks, these techniques raised accuracy from 32% to over 90% against a leading coding agent while reducing cost and latency; specialized search alone improved table-discovery performance by up to 40%, while parallel thinking added latency and token costs before further optimization.


### [Energy trading analytics in a real-time market](https://yomu.fyi/post/energy-trading-analytics-in-a-real-time-market.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Caitlin Gordon
- Published: May 8, 2026

Energy trading markets settle every 15 minutes, while many analytics environments still rely on nightly batch pipelines, creating latency between changing conditions and trading decisions. The post presents Databricks Genie as a conversational interface that lets traders and portfolio managers query actual market data, weather integrations, historical position data, and current books. Its example combines a node’s 90-day natural-gas basis distribution with a Northeast temperature anomaly, while the described system can query real-time and historical data together. Genie also includes settlements, curtailments, and ancillary-service data, with governed access so users see only authorized information. The conclusion is that faster data access can remove information latency and compound margin without replacing trading expertise; the source states that Genie is available today.


### [First-party audience data is the ad sales relationship now](https://yomu.fyi/post/first-party-audience-data-is-the-ad-sales-relationship-now.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Elena Tesser
- Published: May 8, 2026

The post argues that media companies face a stronger need to turn first-party audience data into an advertising sales advantage as third-party cookies decline and buyer attribution expectations rise. Advertisers are buying audiences rather than inventory, so sales teams need granular audience insights, validated campaign performance, clear attribution, and revenue context to compete in RFPs. Databricks Genie is presented as a natural-language interface to governed first-party data, allowing leaders to query defined audience segments such as users who completed at least 75% of automotive content in 90 days, segmented by income quartile and household ownership. The same analytical environment connects demographic composition, content and behavioral signals, post-campaign measurement, attribution, and revenue pacing against budget and prior year.


### [Operating room utilization is hiding in your scheduling data](https://yomu.fyi/post/operating-room-utilization-is-hiding-in-your-scheduling-data.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Adam Crown
- Published: May 8, 2026

Operating room utilization measures in-room surgical minutes against allocated block-time minutes, yet most US health systems reportedly run at 65–75% versus an 80% industry target. Daily performance reports arrive the next morning, after schedules are set and opportunities to release unused blocks, redeploy staff, or backfill add-on cases may have passed. The post presents Databricks Genie as a natural-language interface for querying scheduling, utilization, and outcomes data without a data analyst request. Its proposed analytical environment combines scheduling data, actual case logs, block-release records, contribution margin, and staffing costs, with breakdowns by surgeon, service line, facility, and day of week. Genie surfaces specific intervention targets, although the post says it provides data access rather than automating OR management.


### [Why telecom churn prediction misses the intervention window](https://yomu.fyi/post/why-telecom-churn-prediction-misses-the-intervention-window.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Elena Tesser
- Published: May 8, 2026

Telecom churn programs often intervene after customers have already shifted behavior, contacted support, or decided to leave, even though earlier signals exist in operational data. The post frames the gap as organizational: propensity models may be sophisticated, but retention leaders need timely, specific answers about high-value customers, likely triggers, and historically effective interventions. Databricks Genie is presented as a natural-language interface over customer behavioral and commercial data, able to surface targets such as premium postpaid customers with usage declines above 20%, recent support contacts, and contracts ending within 90 days. Its described capabilities combine usage, support, billing, network experience, competitive tenure, and intervention history while supporting segment and individual analysis. The proposed operating model prioritizes interventions by customer lifetime value and aims to act early enough for retention efforts to change outcomes.


### [Growth analytics is what comes after growth hacking](https://yomu.fyi/post/growth-analytics-is-what-comes-after-growth-hacking.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Madelyn Mullen
- Published: May 8, 2026

Growth analytics is presented as the next stage after growth hacking, arguing that competitive acquisition now depends on precise unit economics, cohort quality, and rapid conversion optimization rather than isolated tactics. It distinguishes product analytics, which explains feature use and user flows, from growth analytics, which connects acquisition sources and costs with activation, revenue, retention, and cohort behavior. The stated bottleneck is fragmented tooling: answering questions such as 90-day LTV by channel correlated with seven-day activation can require manually joining multiple systems and delay weekly budget decisions. Databricks Genie is described as a governed natural-language interface over unified acquisition, behavioral, and billing data, supporting attribution-to-LTV analysis, payback modeling, and paid-versus-organic comparisons. The text reports a 50% relative acquisition-rate lift, from 8% to 12%, for customers using Genie, and shorter onboarding insight cycles from months to weeks.


### [Real-world evidence for medical affairs: who can actually use it?](https://yomu.fyi/post/real-world-evidence-for-medical-affairs-who-can-actually-use-it.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Adam Crown
- Published: May 8, 2026

Real-World Evidence (RWE) is increasingly sought by payers, providers, and regulators, but the post distinguishes it from abundant Real-World Data (RWD), which must undergo rigorous study design, analysis, and interpretation to become credible evidence. It identifies four Medical Affairs use cases: regulatory submissions and commitments, payer formulary discussions, rapid HCP scientific exchange, and internal pipeline or portfolio decisions. The stated operational problem is that teams with claims, EHR, registry, and other assets often lack the fluency or capacity to answer complex questions within competitive and regulatory timelines. Databricks Genie is presented as a natural-language interface that can query unified RWE assets, including a treatment-initiation and 12-month-persistence example that reportedly surfaces in seconds rather than requiring several days of data-science work. The post also describes governance, logged and attributable requests, treatment-pathway awareness, and access controls for MSL support.


### [Wealth advisor productivity starts with the client conversation](https://yomu.fyi/post/wealth-advisor-productivity-starts-with-the-client-conversation.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Kim Hatton
- Published: May 8, 2026

Wealth management client reviews are intended to turn portfolio information into advisory conversations, but substantial meeting time is spent confirming allocations, benchmark performance, and major positions. Preparing for these meetings requires advisors to combine portfolio, tax, estate, and goal-related information from multiple systems for each client, often under time pressure. Databricks Genie is presented as a conversational way to query client portfolio data in real time, including a tax-loss-harvesting question that combines unrealized losses with an estimated current-year tax rate. Its described capabilities include unified holdings, transaction, tax-lot, and profile data; household-level analysis; authorization controls; and external market context, with the stated aim of reducing information preparation so advisors can conduct deeper, more personalized conversations.


### [How lakebase architecture delivers 5x faster Postgres writes](https://yomu.fyi/post/how-lakebase-architecture-delivers-5x-faster-postgres-writes.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: David Wein, Vlad Lazar
- Published: May 7, 2026

Lakebase’s separation of Postgres compute and distributed storage removes the local-disk torn-page risk that necessitates full page writes (FPW) in traditional deployments. The change disables FPW on compute and moves full page image generation into the pageserver, which creates reset points after a configured number of WAL deltas while preserving bounded read reconstruction. In HammerDB TPROC-C tests, 32-vCPU throughput rose from 95,686 to 439,300 new orders per minute, while average WAL per transaction fell from 58KB to under 4KB, a 94% reduction. Production validation reported WAL reductions from 30MB/s to 1MB/s on a 56-vCPU project, p99 read latency improvements of 30% to 50%, and up to threefold regional storage p99 improvement. The rollout reached all Lakebase Serverless and Neon databases without restarts or customer interruption, using the existing XLOG\_FPW\_CHANGE mechanism.


### [Why talent transformation is the missing focus of enterprise AI](https://yomu.fyi/post/why-talent-transformation-is-the-missing-focus-of-enterprise-ai.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Christy Seto, Pratyarth Rao
- Published: May 7, 2026

Enterprise AI adoption is presented as a talent problem as much as a technology or strategy problem: insufficient worker skills are identified as the biggest barrier to integrating AI into work, while hiring alone cannot fill demand. The proposed response is continuous upskilling for both technical practitioners and line-of-business users, replacing one-time enablement with learning tied to current work and changing platform capabilities. Databricks Academy Pro is an annual, per-seat subscription combining self-paced courses, weekly live reviews, unlimited public instructor-led classes, hands-on labs in hosted Databricks environments, and one certification exam voucher per user each year. Starter, Growth, and Enterprise tiers add private online community access, a custom Academy portal, and a designated Talent Transformation program manager as seats increase. The offering is available today, alongside private instructor-led training and standalone certification vouchers.


### [Public health intelligence shouldn't require a data scientist](https://yomu.fyi/post/public-health-intelligence-shouldn-t-require-a-data-scientist.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Kacey Hertan
- Published: May 7, 2026

State, tribal, local, and territorial (STLT) health agencies manage data across surveillance, vital records, Medicaid, WIC, and emergency preparedness systems, but those systems are fragmented and difficult to query together. That limits rapid public health intelligence: questions linking emergency-department visits with pharmacy dispensing, school absenteeism, vaccination, demographic, or geographic data can require epidemiologists to assemble manual queries over weeks, even when decisions require answers within hours. The post presents Databricks Genie as a natural-language interface for querying this environment, backed by a Databricks engine that handles petabyte-scale datasets across real-time streams and historical records. It describes cross-program synthesis, Unity Catalog row- and column-level access controls, HIPAA-compliant governance, traceability to the underlying query, and validation controlled by health experts. Examples include county-level influenza-like illness trends overlaid with vaccination coverage and identifying counties with high opioid overdose rates and low treatment utilization; Genie is described as available today.


### [Mean time to detect is a data access problem](https://yomu.fyi/post/mean-time-to-detect-is-a-data-access-problem.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Taylor Kain
- Published: May 7, 2026

Security operations centers measure MTTD, MTTR, false-positive rates, and analyst utilization, yet investigations often stall because analysts assemble evidence across fragmented systems. A single alert may require separate queries for logs, identity records, asset information, prior alerts, and cross-source timelines, making the analyst the integration layer and creating an MTTI bottleneck. The post presents Lakewatch with Databricks Genie as an agentic interface powered by Anthropic Claude models: analysts ask natural-language questions while autonomous agents hunt, summarize, correlate, and reconstruct timelines across security, IT, and business data. It argues that this architecture can reduce investigation work from manual, multi-system workflows to answers in seconds, while retaining analyst-level access controls and governed data access as exploit time has shrunk to 1.3 days.


### [Rethinking Distributed Systems for Serverless Performance and Reliability](https://yomu.fyi/post/rethinking-distributed-systems-for-serverless-performance-and-reliabil.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Aaron Davidson, Roland Fäustlin, Zach Williams
- Published: May 6, 2026

Building serverless compute for Apache Spark requires more than warm machine pools or basic autoscaling because traditional deployments couple user applications to infrastructure, allowing contention and inefficiencies to undermine performance and reliability. The proposed architecture separates these concerns through Spark Connect’s client-server model over gRPC, a gateway that routes workloads using query size, cluster utilization, and latency profile, and an adaptive autoscaler that adjusts capacity horizontally and vertically. Spark Connect isolates user applications from drivers, while the gateway continually re-evaluates placement to reduce interference between workloads. The autoscaler offers Standard and Performance-Optimized modes and can respond to out-of-memory errors by restarting tasks on larger VMs without manual intervention. Reported outcomes include a 99.998% upgrade success rate across more than 4.5 billion workloads, 2–5x faster Unilever pipelines, and operational-cost reductions of 25%.


### [Responsible AI Governance: A Practical Framework for Business Leaders](https://yomu.fyi/post/responsible-ai-governance-a-practical-framework-for-business-leaders.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: May 6, 2026

Responsible AI governance is presented as an operational framework for leaders overseeing systems that can produce biased outputs, expose sensitive data, and create regulatory, financial, or reputational harm. It draws on the NIST AI RMF and OECD AI principles, maps to EU AI Act requirements, and uses human dignity, fairness, privacy, accountability, transparency, and security as governance values. The program starts with a living inventory recording purpose, ownership, training-data sources, affected populations, review dates, model lineage, and third-party status, followed by risk classification and assessments based on potential impact. It calls for lifecycle controls including bias mitigation, security testing, human review, drift monitoring, audits, incident exercises, and confidential concern reporting. The roadmap recommends piloting governance on a highest-risk product line, scaling controls across business units, and reviewing the framework annually or after major incidents, regulatory updates, or portfolio changes.


### [Machine Learning Use Cases: Practical Industry Applications](https://yomu.fyi/post/machine-learning-use-cases-practical-industry-applications.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: May 6, 2026

This guide surveys practical machine learning applications across industries, from medical imaging and fraud detection to demand forecasting, document processing, and customer-service automation. It defines supervised, unsupervised, semi-supervised, and reinforcement learning, then relates technique selection to the business question, data type, and availability of labels. Technical coverage includes convolutional neural networks for image analysis, transformers and large language models for generative AI, and time-series workflows using cleaning, lag features, and forecasting. The operational guidance addresses feature stores, experiment tracking, CI/CD, drift monitoring, retraining, cost optimization, fairness, privacy, explainability, and model risk management. Templates and checklists structure projects around business problems, data sources, metrics, architecture, measured outcomes, validation, governance, and escalation, while cited Databricks case studies and tools provide implementation examples.


### [The AI scaling gap hiding in digital native companies](https://yomu.fyi/post/the-ai-scaling-gap-hiding-in-digital-native-companies.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Madelyn Mullen
- Published: May 5, 2026

An Economist survey of 1,220+ executives across eight industries finds that digital native companies lead in ambition and breadth of AI deployment, but not in full operational maturity. At 18%, they are the most likely group to prioritize embedding AI across core processes at scale, and nearly 92% report AI ROI ahead of plan. Yet they lead fully embedded AI—defined as use by 100+ users, SLAs, and performance and impact monitoring—in only one of eight functions, R&D/product development; they rank seventh in finance and sixth in operations and supply chain. Telecom, media and entertainment, manufacturing, and energy outperform them in selected functions despite lower stated scaling priority. The proposed response is to build shared, production-grade foundations for data, governance, workloads, models, agents, and applications, with security, lineage, monitoring, and performance measurement treated as reusable capabilities.


### [10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks](https://yomu.fyi/post/10-trillion-samples-a-day-scaling-beyond-traditional-monitoring-infra.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: David Yuan, Yi Jin, Karan Bavishi, HC Zhu, Joey Beyda
- Published: May 5, 2026

Databricks’ monitoring infrastructure now tracks 5 billion active timeseries in real time and ingests more than 10 trillion samples daily, exposing scalability, reliability, cost, and operability limits in its older stack. The replacement, Pantheon, is a fork of CNCF Thanos deployed across more than 160 instances in about 70 regions and three cloud providers; tiered storage, differentiated memory retention, isolated replicated Receive groups, multitenancy, and a custom control plane support automated scaling and recovery. Pantheon’s largest instance holds about 300 million in-memory timeseries and handles nearly 1,000 PromQL queries per second, while migration reduced annual cloud costs by millions and monitoring downtime by roughly five times. For high-cardinality troubleshooting, Hydra preserves raw metrics in Delta tables, exposes them through Grafana and SQL, and unifies metric semantics across aggregated and raw paths, with freshness improvements planned.


### [AI success starts with clean data, not just better models](https://yomu.fyi/post/ai-success-starts-with-clean-data-not-just-better-models.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Aly McGue
- Published: May 5, 2026

Kraken’s data transformation work argues that successful AI depends on clean, unified, accessible data rather than model quality alone. Serving more than 90 million customer accounts across 27 countries, the platform uses Databricks to distribute data securely and at scale, while clients need documentation, join logic and business context to make it useful. Unification reduces the analyst bottleneck, builds trust in shared numbers and enables self-service analytics, including conversational querying through Databricks Genie. The discussion also describes metadata as a live model input: Unity Catalog and Delta Sharing let Kraken share context alongside data instead of relegating it to PDFs or separate web pages. Reported client examples include call-center dashboards updated every few hours with predictive models and faster tariff experimentation, while organizations with stronger data skills and culture are positioned to adopt agentic AI more quickly.


[Newer posts](https://yomu.fyi/company/databricks/page/16.md) · [Older posts](https://yomu.fyi/company/databricks/page/18.md)
