Loading…

Databricks
Data and AI platform for data engineering, analytics, machine learning, and generative AI.
Latest articles
What’s new in Genie Code at Data + AI Summit 2026
At Data + AI Summit 2026, Databricks announced expansions to Genie Code for complex, agentic data and ML work. The changes include a full-page command center for managing concurrent threads and assets, upgrades across production ML engineering, and scheduled tasks that run prompts while users are away. For ML workflows, Genie Code uses Databricks production expertise and Genie Ontology, integrates with MLflow and Model Serving, and can move GPU jobs to AI Runtime while using workspace environment features. It can write features, coordinate edits, run and debug code, compare candidates, inspect endpoint health, and diagnose issues, with users deciding what to keep. Scheduled tasks are described as coming soon, creating reviewable threads from prompts and optional Databricks assets.
Julia Powell, Gal Oshri, Weston HutchinsWhat’s new in Databricks Data + AI Platform security and compliance at Data + AI Summit 2026
At Data + AI Summit 2026, Databricks announced security and compliance capabilities for scaling Genie, Lakebase, serverless workloads, and AI-powered applications without relying solely on manual provisioning, static network controls, or siloed compliance programs. Automatic Identity Management (AIM) for Microsoft Entra ID is generally available on AWS and Google Cloud, AIM for Okta is in Public Preview, and Context-Based Ingress is in Public Preview across all three clouds for policies based on network source, identity, and access scope. Private Network Gateway, in Private Preview on Azure Databricks, provides one secure connection from serverless workloads to private networks, while expanded Private Link support extends to Lakebase and other services. Compliance additions include Azure Serverless coverage, HITRUST across AWS, Azure, and Google Cloud, ISMAP on Azure and AWS, expanded AWS GovCloud availability, and planned FedRAMP High support on Azure Commercial.
Jason Wu, Samrat Ray, Filippo Seracini, Alex Esibov, Vijay Raja, Kelly Albano, Robert Zhang, Mia Penfold LopezBuilding an open ecosystem for AI governance with Unity AI Gateway
Databricks announced the Unity AI Gateway partner ecosystem, extending enterprise AI governance beyond models to runtime interactions among models, agents, MCP servers, skills, and AI tools. Built on Unity Catalog, the gateway lets organizations apply policies, monitor activity, manage spend, and govern AI across providers and frameworks, while integrating security, identity, and governance products they already use. The announcement groups the integrations into runtime AI security, observability and guardrails; agent identity and access governance; and AI observability and risk monitoring. Named integrations include Alice, CrowdStrike Falcon AI Detection and Response, Cyera, HiddenLayer, Netskope, Noma Security, Obsidian Security, Openlayer, Okta, Ping Identity, SailPoint, and Saviynt, with described capabilities including prompt-injection detection, data-loss prevention, agent discovery, authorization, and lifecycle governance.
David Nasi, Kelly Albano, Ashish KathapurkarWhat’s New in the AI Platform: Agents for ML Engineering, Our Deep Learning Platform, and New Capabilities for Real-Time ML
The announcement presents three additions to the Databricks AI Platform: Genie Code support for ML engineering, AI Runtime’s serverless GPU environment, and expanded real-time ML capabilities. Genie Code integrates with Unity Catalog, Feature Store, training, serving, monitoring, and MLflow, assisting with feature engineering, model training, deployment, evaluation, and production operations. AI Runtime provides on-demand serverless NVIDIA A10 and H100 GPUs, supports high-performance multinode training with RDMA and high-performance data loading, and adds Lakeflow Jobs, DABs, MLflow, and Unity Catalog integration. For real-time ML, the platform adds declarative feature engineering, streaming features, online feature serving on Lakebase, and enhanced Model Serving targeting 300K+ QPS with under 10ms p99 latency overhead. Reported customer examples include faster workflows, lower infrastructure costs, and production scaling beyond 100K QPS.
Tejas Sundaresan, Mike Del BalsoIntroducing the Agentic CDP: A New Species of CDP for a New Era of Agents
Traditional customer data platforms (CDPs) were built for human-managed, batch-based campaigns, but the post argues that agentic buying requires millisecond speed, hyper-personalization, and richer context. It contrasts the familiar Golden Record with Golden Context, which combines customer data with current business goals and the history and outcomes of prior decisions. The proposed Agentic CDP uses “Infinity Campaigns,” always-on engagement loops that use LLMs and agents to adapt messaging, timing, and channels for individuals. It is also embedded in the data foundation, bringing customer, business, and decision context together under existing governance, and is designed for agents and humans from the outset. Databricks presents CustomerLake as an implementation of these principles for its platform.
Tasso Argyros, Ali Ghodsi, Reynold XinWhat is data pipeline architecture?
Data pipeline architecture is the end-to-end blueprint for collecting, processing, storing and delivering data from source systems to people, applications and models. It distinguishes logical design, which defines stages and responsibilities, from physical design, which assigns tools and infrastructure to those stages. The common four layers are ingestion, processing and transformation, storage, and serving and consumption, with orchestration and observability spanning the pipeline. The brief compares batch, streaming, Lambda, Kappa and medallion patterns, explaining trade-offs involving freshness, cost, complexity and operational burden. It also contrasts ETL with ELT and presents governance, monitoring and right-sized processing as reliability practices, concluding that architecture should match the use case and balance freshness, cost and reliability.
Databricks StaffEnabling Governed Vibe Coding for Enterprise Apps on Databricks
Databricks introduces three capabilities intended to bring vibe coding to enterprise applications, where speed alone does not provide business-data context, deployment safety, or cost control. App Spaces lets admins define resource and data access, on-behalf-of-user API scopes, and security policies for groups of apps, with each app inheriting those settings. Genie App Builder turns plain-language descriptions into working internal apps through generated plans, live previews, AppKit, and awareness of workspace data assets and Unity Catalog semantics. Serverless micro apps run in isolated lightweight virtual machines, start quickly when needed, scale to zero while idle, and use usage-based rather than reserved-capacity infrastructure. Together, the capabilities are presented as a way for business-proximate users to build on enterprise data while organizations apply consistent governance and support broader app portfolios; all three are coming to Databricks Apps, with private previews coming soon.
Evan Pandya, Justin DeBrabant, Cong XuIntroducing OpenSharing: the Next Evolution of Delta Sharing for the Agentic Era
OpenSharing is presented as the next evolution of Delta Sharing, extending an open zero-copy data-sharing protocol from tables and files to models, agents, semantic context, unstructured data, and reusable AI logic. The protocol is now an independent open-source project hosted by the Linux Foundation, while Databricks OpenSharing adds Unity Catalog governance and audit logging, Marketplace discoverability, and enterprise features. Genie Agent Sharing supports governed AI experiences across organizational boundaries, with controls for proprietary instructions, data access, daily prompt quotas, and row exports. SecureConnect removes per-recipient firewall changes through a Databricks-managed proxy, while Global Distribution uses local replicas to reduce egress fees and latency. The launch also supports Apache Iceberg REST Catalog API, external catalogs, and on-premises storage partners; providers define shares in Unity Catalog, recipients query live data through existing tools, and governance enforces access controls.
Huey Han, Harish Gaur, Akram Chetibi, Mengxi ChenAnnouncing Apps on Databricks Marketplace
Databricks announces the Public Preview of Apps on Databricks Marketplace, allowing customers to discover, install, and run third-party data and AI applications inside secure Databricks workspaces. The offering addresses procurement challenges involving data movement, lengthy security reviews, custom integrations, and fragmented identity management by bringing applications to the customer’s data. Installed apps run in isolated sandboxes within the consumer’s Databricks account, inherit Unity Catalog governance, and use dedicated serverless compute with consumer-controlled external access through Serverless Egress Gateway policies. Providers can publish closed-source containerized apps once for no-egress distribution without maintaining per-customer infrastructure, while applications connect natively to services including SQL Warehouse, Lakebase, Model Serving, and Foundation Model APIs. The Public Preview launches with 20 partners, and planned additions include bundled assets, provider analytics, and commercial monetization.
Tia Chang, Akram Chetibi, Harish Gaur, Stephen Orban, Mengxi ChenIntroducing OpenSharing SecureConnect
OpenSharing SecureConnect addresses the networking burden of sharing live data from provider storage behind private networks, where providers and recipients otherwise exchange firewall and egress details manually. It is a Databricks-managed proxy that routes recipient storage access through Databricks endpoints after a one-time provider setup, while the data remains in the provider’s bucket. Providers allowlist Databricks Serverless Data Plane endpoints and enable SecureConnect for a metastore; serverless recipients require no configuration, while classic and open recipients allowlist stable inbound IPs. Optional NCC provides private link connectivity, and mutual TLS is available for recipients. SecureConnect is in Public Preview and supports cross-region, cross-cloud sharing plus customer-managed and Databricks Default Storage.
Huey Han, William Chau, Harish GaurSciene AI Companion: building an autonomous Customer Success platform on Databricks
Sciene built AI Companion for Quartile’s Customer Success organization, where CSMs support more than 1,000 brands and previously spent substantial time preparing decks, reconstructing context, and investigating account changes. The platform addresses personalization at scale, high-volume content generation, and root-cause diagnosis by combining account data, CSM communication styles, company principles, and cross-domain business data. Its Email Hub cuts reply time from 15–30 minutes to about three minutes, Meeting Hub reduces preparation for 80+ slide decks from over two hours to around 10 minutes, and Account Flagging reduces diagnosis of flagged accounts from 30+ minutes to about five. Databricks provides the shared governed foundation: Delta Sharing supplies data without copies, Lakebase stores operational state, and SQL Warehouses serve analytical, AI, and operational workloads from the same tables. The design keeps CSMs responsible for judgment while giving them current context for customer interactions.
Renata Fencz, Solano Campos, Rodrigo Mohr, Ricardo MorandiniWhat’s new with Unity Catalog at Data + AI Summit 2026
At Data + AI Summit 2026, Unity Catalog announcements position the catalog as a runtime governance layer for enterprise data and AI, organized around control, context, and choice. Control additions include Unity AI Gateway for governing models, agents, MCP services, skills, and tools; contextual service policies can allow, deny, or require approval for runtime actions, while budgets, hard caps, tracing, and guardrails address spend, investigation, and safety. Context additions include Glossary and Domains for business meaning and scoped asset organization, plus Metrics that standardize KPIs for SQL, BI tools, APIs, and agents; Genie Ontology is described as a continuously learned enterprise context layer. Choice additions span cross-cloud and cross-region addressability, managed disaster recovery, Delta and Iceberg interoperability, multimodal and geospatial types, and open sharing of data, AI assets, and applications across organizations.
The Unity Catalog Product and Engineering TeamAgent Bricks: Data + AI Summit 2026
At Data + AI Summit 2026, Databricks announced Agent Bricks as a comprehensive developer platform for building and operating agents, extending a product launched the previous year. The announcement frames the core agent loop as only 1% of the work, with token capacity, deployment, security, evaluation, monitoring, context, and sharing forming the remaining infrastructure burden. Agent Bricks addresses choice, context, and control through support for multiple proprietary, open-source, and custom models, any agent harness, MCP-connected data, Genie Ontology, managed memory, document intelligence, sandboxes, and governed tools. Unity AI Gateway adds catalogs, fine-grained access controls, budgets, traffic routing, contextual policies, monitoring, and registry support for agents, tools, and models. Databricks says more than 100,000 agents have been built and customers including AstraZeneca, 7-Eleven, Fox Corporation, and Block have shipped agents on the platform.
Hanlin Tang, Kasey Uhlenhuth, Akhil Gupta, Patrick WendellAI governance at Data + AI Summit 2026: What’s new with Unity AI Gateway
Databricks announces new Unity AI Gateway capabilities for governing enterprise AI as organizations operate multi-model, multi-agent, and multi-vendor estates connected to models, MCP services, APIs, and tools. The update adds unified spend visibility, granular attribution, hard spend caps, and smart routing, alongside Unity Catalog support for registering and governing models, MCP services, agents, and skills. Contextual Service Policies, in Beta, can allow, deny, or require approval for actions based on users, agents, models, tools, services, or request and response contents, with guardrails for risks such as PII exposure and prompt injection. The announcement also covers end-to-end tracing, coding-agent analysis with Genie, incident investigation with Lakewatch, ecosystem integrations, and Managed Omnigent on Databricks in Beta.
David Nasi, Stefania Leone, Ahmed Bilal, Kevin Stumpf, Martin Grund, Vladimir Kolovski, Kelly AlbanoLakeflow: A new era of agentic data engineering
Databricks announces a major evolution of Lakeflow, its unified platform for data engineering across ingestion, transformation, and orchestration, with capabilities centrally governed by Unity Catalog. Genie Code and generally available Lakeflow Designer support agentic and no-code pipeline development, while Genie ZeroOps monitors production assets, analyzes failures, proposes fixes, and validates them in a governed sandbox before human approval. Lakeflow Connect expands to more than 100 managed connectors, and Zerobus Ingest adds Kafka-compatible, gRPC, REST, SDK, and OpenTelemetry interfaces for high-volume event ingestion. Real-Time Mode for Spark Declarative Pipelines reaches Public Preview with end-to-end latency as low as 5 milliseconds, alongside declarative APIs and expanded Lakeflow Jobs integrations. The release also adds data-readiness triggers and external orchestration for systems including Snowflake, REST APIs, Slack, and PagerDuty.
Bilal Aslam, Ray Zhu, Manish Dalwadi, Saad Ansari, Giselle GoicocheaIntroducing Genie ZeroOps: Put your data and AI operations on autopilot
Genie ZeroOps is an autonomous background agent for monitoring and operating data and AI assets, including jobs, pipelines, tables, and ML models. It continuously detects visible and silent failures, uses Unity Catalog lineage and platform observability to assess root causes, generates remediation through development workflows, and verifies fixes in isolated sandboxes. These environments use shallow, zero-copy table clones, scoped permissions, and network isolation, so proposed changes run against real data without touching production or applying anything before approval. For ML workloads, the agent can diagnose degraded predictions, train a candidate on corrected features, evaluate it against the production model’s existing eval suite and criteria, and support live-traffic ramping when it is measurably better. Genie ZeroOps is entering private preview in the coming weeks, initially supporting jobs, pipelines, tables, and ML workloads; Apps and Lakebase databases are on the roadmap.
Bilal Aslam, Lennart Kats, Ray Zhu, Mike Del Balso, Ori ZoharAnnouncing Lakebase Search: agent-native retrieval built into Lakebase Postgres
Lakebase Search is a beta offering on AWS and Azure that adds hybrid vector and full-text retrieval to Lakebase Postgres. It uses the lakebase_vector and lakebase_text extensions to keep retrieval, memory, operational data, and hybrid search in one backend. lakebase_vector retains pgvector types and operators, applies RaBitQ clustering and compression for 32x smaller indexes, and targets more than 1B vectors, while lakebase_text replaces GIN with object-storage-optimized BM25 ranking. A tiered cache keeps hot data on NVMe and places colder data in object storage; the source reports lower memory needs, faster index builds, and cold-cache startup than standard pgvector HNSW in its LAION-100M benchmark. The extensions also combine vector similarity and keyword relevance with reciprocal rank fusion in a single SQL query, enabling joins and tenant filtering alongside transactional workflows.
Pranav Aurora, Zhou Sun, Jinjing ZhouAnnouncing the new Databricks Startup Program
The Databricks Startup Program has been updated for venture-backed, early-stage startups building with data and AI. Qualifying companies can receive up to $200,000 in credits across Databricks and Neon, along with hands-on technical guidance, partner and community access, and connections to go-to-market teams and founder communities. The program is aimed especially at startups that recently raised institutional funding from pre-seed through Series A, and it is intended to provide an app backend, data, and AI stack from idea through product-market fit. Databricks and Neon are presented as providing a database and access to foundation models on day one, while Databricks supplies analytics, data warehousing, AI systems, and enterprise-grade governance as companies launch. Applications are available through the Databricks Startup Program.
Brad Van Vugt, Arjun RajeswaranUnifying Data and Governance in the Agentic Era: What’s New with Azure Databricks
At Data + AI Summit 2026, Azure Databricks announced capabilities aimed at moving enterprises from experimental AI pilots to production-grade automated workflows by unifying data, productivity tools, marketing, and governance on Azure. Its Agentic Data foundation introduces LTAP, combining analytical data, streaming pipelines, and live application transactions in one lakehouse storage copy; Lakebase adds a managed serverless Postgres engine with copy-on-write branching, while Lakehouse//RT targets millisecond responses for high-concurrency workloads. Genie integrations for Microsoft Teams, M365 Copilot, Excel, and SharePoint bring governed lakehouse intelligence and ingestion into daily work, alongside tools for agents, applications, pipelines, and autonomous operations. CustomerLake adds Profile Agents and Campaign Agents for customer profiles and personalization, while Genie Ontology and Unity AI Gateway provide semantic context, rate limits, content filtering, and spend controls.
Isaac Gritz, Toussaint Webb, Ben Tripp, Kiriana StukasIntroducing Genie One, Genie Agents, and Genie Ontology
Databricks announces Genie One, Genie Agents, and Genie Ontology to help enterprises answer business questions and act on data whose context is scattered across dashboards, queries, documents, tickets, and chats. Genie One connects data and business tools through Lakehouse federation, Lakeflow Connect, native integrations, Slack, Teams, mobile apps, schedules, alerts, document creation, custom skills, and MCP support. Genie Agents evolve Genie Spaces into domain-specific agents that can reason over structured and unstructured data, execute multi-step workflows, and be created from a prompt. Genie Ontology builds a permission-aware living graph from enterprise assets, weighting sources by authority, usage, certification, and freshness. In an internal 28-question benchmark, Genie answered 84.5% correctly on the first attempt and delivered twice the speed of the strongest coding agent.
Sydney Sundell, Ken Wong, Elise Georis