Loading…

Databricks
Data and AI platform for data engineering, analytics, machine learning, and generative AI.
Latest articles
Azure Databricks at Data + AI Summit 2026 featuring Industry Leaders and Partners
Data + AI Summit 2026 brings Databricks and Microsoft leaders, partners, and customers together June 15–18, with in-person and virtual programming focused on Azure Databricks. The collaboration presents Azure Databricks as a first-party Microsoft offering for unifying data, analytics, and AI on a secure, scalable foundation, with sessions covering ecosystem integration, federated analytics, governance, modernization, and AI applications. One technical example introduces zero-copy federation between Azure Data Manager for Energy and Databricks compute, preserving ADME as the source of truth while avoiding large-scale data copies. Another shows Unity Catalog External Locations extending governed access to Microsoft OneLake without ETL pipelines, while customer sessions describe Apache Iceberg and Apache Spark integration, fragmented data consolidation, and production-grade finance workflows using Azure Document Intelligence.
Kiriana StukasEmpower your healthcare agents with ready-to-use MCP on Databricks Marketplace
Databricks announces ready-to-use Model Context Protocol (MCP) servers for healthcare and life sciences through Databricks Marketplace, addressing the need to combine curated biomedical knowledge, timely data, specialized tools, and private records. Listings include services for drug and target intelligence, literature, clinical trials, FDA information, Medicare coverage, ontologies, real-world evidence, clinical semantics, and interoperability, while Climb connects live public sources with private Gold-layer data under Unity Catalog governance. Marketplace and custom MCP servers are centralized in the MCP Catalog and governed by Unity AI Gateway, with Genie Spaces, AI Search, Unity Catalog functions, and SQL Warehouses available as managed MCP servers. Users can assemble agents in AI Playground, Agent Bricks, or notebooks, then deploy endpoints or apps with MLflow tracing, evaluation, human feedback, and AI guardrails; examples span molecular-property lookup, clinical questions, and drug research.
Yen Low, Mark Lee, Matthew Giglia, Nicholas Siebenlist, Jay Bhankharia, Paul Ford, Itai WeissHow Ecolab rebuilt retail intelligence on Databricks and Anthropic Claude
Ecolab needed to combine audits, health inspections, pest telemetry, and other data from nine systems so retail teams could answer location-specific compliance questions. Its Retail Intelligence application is a native Databricks App using Lakebase Postgres, Lakeflow, and Spark Declarative Pipelines to move governed data into a Unity Catalog lakehouse, while Foundation Model APIs serve Claude Sonnet, Claude Haiku, and Gemini. A Coordinator Agent delegates requests to specialized agents that use Vector Search, SQL, Unity Catalog Functions, and an external MCP server; a Response Agent returns cited answers, with short- and long-term memory stored through Lakebase. The system also applies five Judge LLMs, MLflow tracing, and ai_query() batch inference. Report preparation fell from two weeks to under two minutes, while the assistant supports approximately twelve languages at about 98% accuracy.
Babu Chinnaswamy, Nicholas Dylla, Alissa Ellingson, Harish GaurStop building data products. Start building data services.
Howden’s rapid acquisition pace exposed limits in an enterprise data model built around one product per use case, downstream quality checks, and dashboard-driven consumption. Group Chief Data Officer Barry Panayi describes shifting to open, governed data services, moving mastering and quality checks closer to ingestion, and codifying reconciliation in the Accord data model. On Databricks, the company consolidated more than 100 sources of record, standardized pipelines and shared code, and built reusable assets for cross-domain analytics, while continuing to productionize models as consistent services. The account argues that AI agents require a composable services layer, and that insight lag—the time between data existing and being usable—matters more than freshness; conversational analytics through Genie also reduced dashboard-building work.
Aly McGueScaling AI Through Data Fluency
Aer Lingus is redirecting a significant share of its IT and change spending from traditional maintenance toward a Databricks-powered data foundation, addressing legacy systems that trap information in departmental silos. Dave O’Donovan says the airline spent the past 18 months prioritizing platform development, governance, data quality and data literacy rather than chasing each new AI product. Databricks was selected for a unified lakehouse architecture, with data warehousing, Genie’s plain-English querying and real-time operational data intended to broaden access beyond specialist teams. At Aer Lingus’s Operations Control Center, combining sensor and operational inputs gives teams a fuller real-time view for disruption decisions, while commercial teams use live insights to adjust pricing. The transformation also includes a Data Literacy Academy, a 75/25 capacity split between foundational work and innovation, a 20-person Continuous Improvement team, and experiments with agents for business-case development and CFO review.
Aly McGueAWS and Databricks at Data + AI Summit 2026: Accelerating real-world AI innovation
AWS and Databricks describe their expanded collaboration at Data + AI Summit 2026, where AWS returns as a Legend Sponsor with sessions, demos, customer stories, and industry forums. The partnership centers on generative AI adoption, unified governance, and open data architectures, including an agentic stack that combines Amazon Bedrock, Bedrock AgentCore, Kiro, and the Databricks Data + AI Platform. A featured integration uses a governed MCP connection through Databricks Apps so AgentCore can query Unity Catalog-governed data, ask AI/BI Genie questions, and read low-latency state from Lakebase while honoring existing permissions. AWS will demonstrate these workflows at Booth #100 and present a session on federating Unity Catalog to AWS Glue, alongside customer examples including Mastercard, Talkdesk, nCino, Addepar, and Workday. Attendees can also join technical conversations, receptions, and a 14-day Databricks on AWS Marketplace trial with $400 in usage credits.
Sarah Jack, Taylor HossAnnouncing the Public Preview of Custom URLs
Databricks has announced the public preview of Custom URLs, giving each account a single branded domain that serves its workspaces. Previously, workspace-specific URLs made navigation, sharing, bookmarking, and account-wide features more cumbersome, while users had to log in repeatedly when switching workspaces. Custom URLs use a shared account session to verify access and create workspace sessions seamlessly, while preserving authorization boundaries and keeping existing per-workspace URLs functional. The feature provides a unified Genie entry point, cross-workspace Unity Catalog lineage, and an account-level URL that remains stable during disaster-recovery failover, including for downstream tools using existing connection strings. Activation requires Unified Login; Frontend Private Link workspaces fall back to per-workspace URLs, and account admins can claim a URL, enable it, and optionally turn on automatic redirects.
Gordon Wang, Steve Costa, Ankit MishraHow Rivian drives trusted, AI-powered decisions at the speed of thought with Databricks
Rivian is building electric vehicles and services that require fast, trusted decisions across manufacturing, supply chain, finance, service and operational planning, while business users need reliable metrics and insights. Using Databricks AI/BI, Genie, Unity Catalog metric views, Databricks Apps and AI-assisted engineering, the company is consolidating dashboards, semantic definitions, permissions, sensitive data and AI-powered workflows on one governed foundation. Rivian migrated a massive multi-domain dashboard base in less than six months, is standardizing more than 50 metrics, and worked with Databricks as a design partner on roughly 58 product features. The resulting self-service analytics and operational applications cut supply-chain monitoring time by 60 to 70%, reduce inventory investigations from over 30 minutes to under two, predict stock-out risk more than four days ahead, and reduce some ingestion setup time by more than 60%, supporting AI-powered decisions without competing versions of the truth.
Romit Jadhwani, Saritha Suresh, Miranda Luna, Julia PowellJumpstart your Data Modeling with Databricks Industry Data Models
Databricks is publishing a public library of 40 Lakehouse industry data models designed to provide Silver-layer foundations for analytics and machine learning. Each industry offers a Minimum Viable Model and Expanded Coverage Model derived from the same model.json, with breadth rather than attribute depth distinguishing the scopes. A rules-driven AI agent applies more than 200 structural checks across 14-plus modeling domains, enforcing hierarchy, primary and foreign keys, normalization, division balance, data types, governance tags, and acyclic relationships. The models deploy to Unity Catalog in three physical cataloging styles and include DDL, schemas, metric views, classification tags, ontology, diagrams, and synthetic data with valid references. The airline ECM example contains 19 domains, 420 products, 17,278 attributes, 420 primary keys, and 2,877 foreign keys, while the models remain customizable starting points requiring domain expertise and organizational review.
Amr Ali, Drew Triplett, Franco Patano, Shelley ShafferyAI Serving Platform That Adapts to Your Model
Databricks Custom Model Serving addresses the operational burden of serving custom models, whose resource profiles, traffic patterns, and latency requirements vary widely from small CPU classifiers to large GPU-backed language models. Its fully managed platform packages MLflow models and uses isolated Kubernetes deployments, model-appropriate runtimes, and a short request path to limit interference and per-request overhead. At the center, the AutoPilot Pod Autoscaler combines active concurrency and queue signals for horizontal scaling with CPU, GPU, and memory measurements for model-aware target-concurrency adjustment, allowing one controller to adapt across workloads. Warm pools, provisioned concurrency, and zero-downtime updates address startup and deployment concerns, while reported production results include 90%+ cost savings for some customers, up to 2x improvement in p99 and p50 latency, 100K+ QPS, and 99.99% availability.
Anshul GuptaAnnouncing the Databricks storage ecosystem: Governing the enterprise data estate, wherever it lives
Databricks announces a Software-Defined Storage (SDS) Ecosystem for governing enterprise data across on-premises, private-cloud, and edge environments without requiring migration. It responds to sovereignty and regulatory constraints, data-gravity economics, latency requirements, and the need to unlock AI value in backup, archive, and other previously inaccessible data. Storage partners implement the open-source OpenSharing protocol, expose an endpoint, and connect it to Unity Catalog so Databricks Serverless Compute can access governed data where it resides. The integrations provide a unified catalog and support Serverless Compute, Genie, AgentBricks, and model training with zero data movement or duplication; MinIO is generally available, while Everpure, Qumulo, and VAST Data are in preview stages. Databricks also says commitments are in place from six additional providers and that Volumes APIs are being developed to extend OpenSharing to unstructured files for generative-AI workloads.
Rupal Jain, Denis DubeauModern BSA/AML compliance on Databricks
AML operations are strained by fragmented systems, high false-positive volumes, manual case documentation, and opaque vendor scoring, leaving analysts focused on backlog rather than financial-crime intelligence. The proposed Databricks Data + AI Platform unifies transaction monitoring, KYC, sanctions, case history, and policy data under Unity Catalog, using Lakeflow Connect and a Bronze–Silver–Gold Delta architecture with masking, row-level security, and lineage. MLflow, Model Serving, Lakehouse Monitoring, and inference tables support institution-specific detection models, while Agent Bricks coordinates agents for evidence gathering, recommendations, and SAR drafting with analysts retaining final decisions. The architecture also uses Lakebase for governed operational state and Databricks Apps for analyst and executive experiences. Reported outcomes include a 75% reduction in false positives reaching the analyst queue and compressing three-to-six-hour investigations to minutes.
Kateryna Savchyn, Pavithra Rao, Mimi Park, Emerson BayukClaude Fable 5 is now available on Databricks, fully governed through Unity AI Gateway
Claude Fable 5 is now generally available on Databricks, with rollout across AWS, Azure, and Google Cloud through Unity AI Gateway. The Mythos-class model targets long-running, complex, and ambiguous work, including autonomous enterprise workflows, document question answering, code investigation, and multimodal tasks. In Databricks' OfficeQA Pro benchmark, Fable 5 achieved 57.9% correctness, setting a state of the art; compared with Claude Opus 4.8, it was 20% more accurate and used 12% fewer tool calls, but ran approximately 30% slower and generated 2.5x more output tokens. Unity AI Gateway provides unified API access, fine-grained permissions, Unity Catalog logging, request and tool-call guardrails, and spend controls. Agent Bricks supports domain-specific agents, while Anthropic's policy includes 30-day retention for trust and safety purposes only.
Ahmed Bilal, Ivan Zhou, Yash Oza, Gautam Venkatesh, Alice Li, Harish GaurAnnouncing the winners of the 2026 Databricks Customer Awards
The 2026 Databricks Customer Awards recognize organizations and leaders using the Databricks Data + AI Platform across eight categories and four regions. The announcement names winners including Applied Materials, Virgin Atlantic, Fonterra Co-operative Group, Telefónica | Vivo, Virtue Foundation, Octopus Energy, Axpo, Atlassian, Wassym Bensaid at Rivian and Volkswagen Group Technologies, and Kenan Colson at Lippert. Examples include Applied Materials’ move from a Hadoop-based data lake to a governed lakehouse, with 1,500-plus analysts, more than 100 production machine-learning models and faster pipeline development, while Fonterra centralized data to improve supply-chain and compliance work. Rivian unified vehicle, factory and enterprise data, consolidated legacy systems into Delta and Unity Catalog, and designed for 500–600 petabytes; Lippert deployed AI tools across customer care, finance and HR. The announcement presents data and AI adoption as a cross-functional operating model.
Sara SteffenAnnouncing the 2026 Databricks Customer Awards Industry winners
Databricks announced its 2026 Customer Awards Industry winners, recognizing 10 organizations across financial services, communications, health and life sciences, manufacturing, retail and CPG, energy and utilities, enterprise technology, public sector, digital-native businesses and cybersecurity. The cited work uses data and AI to address industry-specific needs, including SMBC Group’s governed lakehouse for risk and finance, Hospital for Special Surgery’s full-system ingestion strategy, Lumen’s conversational service-operations workflows and Superhuman’s high-volume model serving. Reported results include HSS ingesting more than 40 source systems and creating over 14,500 production tables, Lumen recording 3 million-plus AI-powered diagnostics and 35% ticket deflection, and Superhuman handling peaks above 200,000 queries per second. Adobe is also recognized for applying software engineering practices to cybersecurity detection workflows, reducing false positives and improving development speed.
Michael GriffithsTransforming solar and wind maintenance reports with Genie and AI agents
Plenitude and Databricks built an agent-based system that turns solar and wind plant maintenance PDFs into structured data for cross-plant analysis. Event-driven ingestion uses Databricks Jobs and the ai_parse_document AI Function to extract text, tables, figures, and metadata, then stores page- and object-level JSON records in Delta Lake with coordinates, version history, and links to source reports. A Genie space uses Unity Catalog metadata, knowledge-store instructions, and SQL generation to answer natural-language questions, produce visualizations, and export results, while Agent Bricks can orchestrate multi-step workflows and downstream actions. The design also applies automatic liquid clustering to dynamic queries and row-level security to restrict results by country. The resulting data layer supports historical trends, plant comparisons, recurring-fault analysis, and a foundation for predictive maintenance, although the source frames predictive use as a future improvement.
Maria VallarelliEnterprise Data Strategy Roadmap for Business Outcomes
An enterprise data strategy connects organizational data assets to measurable business outcomes, while fragmented architectures can leave data investments uncoordinated and limit real-time analysis and action. The roadmap starts with purpose, scope, executive sponsorship, measurable objectives, KPI mapping, and use-case prioritization based on business impact, feasibility, time to value, and organizational readiness. It then organizes governance, lifecycle management, data quality, target-state architecture, integration, analytics, team structure, compliance, and measurement as interdependent capabilities, emphasizing owners, stewards, decision rights, executable quality rules, and automated cleansing. Implementation proceeds through a time-boxed cross-functional pilot, documented learnings, and incremental scaling, with steering-committee oversight and governance that evolves through feedback. The text gives indicative timelines of 60 to 90 days for a focused pilot, 12 to 18 months for a foundational platform across multiple business units, and multiple years for a mature data-driven culture.
Databricks StaffWhat is Human-in-the-Loop (HITL)?
Human-in-the-loop (HITL) is an AI and machine learning approach that places people in training, supervision, or decision-making to improve accuracy, safety, and ethical alignment. Its feedback loop can include data labeling, output review, escalation, approval, override, and continuous feedback, with confidence thresholds and risk scoring routing only selected decisions to people. The explainer distinguishes HITL, where review occurs before flagged actions, from human-on-the-loop monitoring and human-over-the-loop governance, and separates HITL from RLHF, a training-specific technique. It describes uses in medical imaging, moderation, autonomous vehicles, financial services, and AI agents handling consequential actions. Databricks Agent Bricks is presented as supporting governed traces and Agent Learning from Human Feedback, including a case where 32 feedback items improved instruction-following from roughly 12% to 80%.
Databricks StaffWhat is Explainable AI (XAI)?
Explainable AI (XAI) comprises techniques that help people understand how AI systems produce specific outputs, particularly when machine-learning and deep-learning models operate as black boxes. It distinguishes intrinsically interpretable models, such as decision trees and linear or logistic regression, from post-hoc methods including SHAP, LIME, counterfactuals, saliency maps and Grad-CAM. A typical workflow selects a model and prediction, applies a method suited to the model and audience, reviews outputs such as feature scores or heatmaps, and uses them to assess accuracy, fairness, reliability and compliance. The article stresses that post-hoc explanations are approximations rather than definitive proof, so teams should validate them with domain expertise and, where appropriate, combine methods; MLflow and Unity Catalog can preserve explanation artifacts, lineage and auditability.
Databricks StaffEnabling Evolutionary Database Development: database branching with Lakebase, continued
This installment revisits Evolutionary Database Design and argues that Databricks Lakebase removes the infrastructure constraints that kept several of its database-change practices aspirational. Lakebase is a managed Postgres database with compute separated from shared durable storage, while copy-on-write branches create a new pointer and divergence marker in roughly one second without copying parent data. That enables per-developer, per-PR, and per-experiment databases, real Postgres test branches instead of mocks or in-memory substitutes, and CI isolation at pull-request granularity. The updated playbook adds idempotent migrations, destructive testing, and database-level A/B prototyping, with Unity Catalog governance inherited by branches and agents receiving branches rather than production access. Jen’s example corrupts production-shaped data to test an inventory-code migration and compares column and lookup-table designs, rejecting the latter because its common read path requires a join.
Pramod Sadalage, Kevin Hartman