Loading…
Unity Catalog
112 posts about Unity Catalog. Every summary links to the original.
AWS and Databricks at Data + AI Summit 2026: Accelerating real-world AI innovation
AWS and Databricks describe their expanded collaboration at Data + AI Summit 2026, where AWS returns as a Legend Sponsor with sessions, demos, customer stories, and industry forums. The partnership centers on generative AI adoption, unified governance, and open data architectures, including an agentic stack that combines Amazon Bedrock, Bedrock AgentCore, Kiro, and the Databricks Data + AI Platform. A featured integration uses a governed MCP connection through Databricks Apps so AgentCore can query Unity Catalog-governed data, ask AI/BI Genie questions, and read low-latency state from Lakebase while honoring existing permissions. AWS will demonstrate these workflows at Booth #100 and present a session on federating Unity Catalog to AWS Glue, alongside customer examples including Mastercard, Talkdesk, nCino, Addepar, and Workday. Attendees can also join technical conversations, receptions, and a 14-day Databricks on AWS Marketplace trial with $400 in usage credits.
Sarah Jack, Taylor HossAnnouncing the Public Preview of Custom URLs
Databricks has announced the public preview of Custom URLs, giving each account a single branded domain that serves its workspaces. Previously, workspace-specific URLs made navigation, sharing, bookmarking, and account-wide features more cumbersome, while users had to log in repeatedly when switching workspaces. Custom URLs use a shared account session to verify access and create workspace sessions seamlessly, while preserving authorization boundaries and keeping existing per-workspace URLs functional. The feature provides a unified Genie entry point, cross-workspace Unity Catalog lineage, and an account-level URL that remains stable during disaster-recovery failover, including for downstream tools using existing connection strings. Activation requires Unified Login; Frontend Private Link workspaces fall back to per-workspace URLs, and account admins can claim a URL, enable it, and optionally turn on automatic redirects.
Gordon Wang, Steve Costa, Ankit MishraHow Rivian drives trusted, AI-powered decisions at the speed of thought with Databricks
Rivian is building electric vehicles and services that require fast, trusted decisions across manufacturing, supply chain, finance, service and operational planning, while business users need reliable metrics and insights. Using Databricks AI/BI, Genie, Unity Catalog metric views, Databricks Apps and AI-assisted engineering, the company is consolidating dashboards, semantic definitions, permissions, sensitive data and AI-powered workflows on one governed foundation. Rivian migrated a massive multi-domain dashboard base in less than six months, is standardizing more than 50 metrics, and worked with Databricks as a design partner on roughly 58 product features. The resulting self-service analytics and operational applications cut supply-chain monitoring time by 60 to 70%, reduce inventory investigations from over 30 minutes to under two, predict stock-out risk more than four days ahead, and reduce some ingestion setup time by more than 60%, supporting AI-powered decisions without competing versions of the truth.
Romit Jadhwani, Saritha Suresh, Miranda Luna, Julia PowellJumpstart your Data Modeling with Databricks Industry Data Models
Databricks is publishing a public library of 40 Lakehouse industry data models designed to provide Silver-layer foundations for analytics and machine learning. Each industry offers a Minimum Viable Model and Expanded Coverage Model derived from the same model.json, with breadth rather than attribute depth distinguishing the scopes. A rules-driven AI agent applies more than 200 structural checks across 14-plus modeling domains, enforcing hierarchy, primary and foreign keys, normalization, division balance, data types, governance tags, and acyclic relationships. The models deploy to Unity Catalog in three physical cataloging styles and include DDL, schemas, metric views, classification tags, ontology, diagrams, and synthetic data with valid references. The airline ECM example contains 19 domains, 420 products, 17,278 attributes, 420 primary keys, and 2,877 foreign keys, while the models remain customizable starting points requiring domain expertise and organizational review.
Amr Ali, Drew Triplett, Franco Patano, Shelley ShafferyAnnouncing the Databricks storage ecosystem: Governing the enterprise data estate, wherever it lives
Databricks announces a Software-Defined Storage (SDS) Ecosystem for governing enterprise data across on-premises, private-cloud, and edge environments without requiring migration. It responds to sovereignty and regulatory constraints, data-gravity economics, latency requirements, and the need to unlock AI value in backup, archive, and other previously inaccessible data. Storage partners implement the open-source OpenSharing protocol, expose an endpoint, and connect it to Unity Catalog so Databricks Serverless Compute can access governed data where it resides. The integrations provide a unified catalog and support Serverless Compute, Genie, AgentBricks, and model training with zero data movement or duplication; MinIO is generally available, while Everpure, Qumulo, and VAST Data are in preview stages. Databricks also says commitments are in place from six additional providers and that Volumes APIs are being developed to extend OpenSharing to unstructured files for generative-AI workloads.
Rupal Jain, Denis DubeauModern BSA/AML compliance on Databricks
AML operations are strained by fragmented systems, high false-positive volumes, manual case documentation, and opaque vendor scoring, leaving analysts focused on backlog rather than financial-crime intelligence. The proposed Databricks Data + AI Platform unifies transaction monitoring, KYC, sanctions, case history, and policy data under Unity Catalog, using Lakeflow Connect and a Bronze–Silver–Gold Delta architecture with masking, row-level security, and lineage. MLflow, Model Serving, Lakehouse Monitoring, and inference tables support institution-specific detection models, while Agent Bricks coordinates agents for evidence gathering, recommendations, and SAR drafting with analysts retaining final decisions. The architecture also uses Lakebase for governed operational state and Databricks Apps for analyst and executive experiences. Reported outcomes include a 75% reduction in false positives reaching the analyst queue and compressing three-to-six-hour investigations to minutes.
Kateryna Savchyn, Pavithra Rao, Mimi Park, Emerson BayukClaude Fable 5 is now available on Databricks, fully governed through Unity AI Gateway
Claude Fable 5 is now generally available on Databricks, with rollout across AWS, Azure, and Google Cloud through Unity AI Gateway. The Mythos-class model targets long-running, complex, and ambiguous work, including autonomous enterprise workflows, document question answering, code investigation, and multimodal tasks. In Databricks' OfficeQA Pro benchmark, Fable 5 achieved 57.9% correctness, setting a state of the art; compared with Claude Opus 4.8, it was 20% more accurate and used 12% fewer tool calls, but ran approximately 30% slower and generated 2.5x more output tokens. Unity AI Gateway provides unified API access, fine-grained permissions, Unity Catalog logging, request and tool-call guardrails, and spend controls. Agent Bricks supports domain-specific agents, while Anthropic's policy includes 30-day retention for trust and safety purposes only.
Ahmed Bilal, Ivan Zhou, Yash Oza, Gautam Venkatesh, Alice Li, Harish GaurAnnouncing the winners of the 2026 Databricks Customer Awards
The 2026 Databricks Customer Awards recognize organizations and leaders using the Databricks Data + AI Platform across eight categories and four regions. The announcement names winners including Applied Materials, Virgin Atlantic, Fonterra Co-operative Group, Telefónica | Vivo, Virtue Foundation, Octopus Energy, Axpo, Atlassian, Wassym Bensaid at Rivian and Volkswagen Group Technologies, and Kenan Colson at Lippert. Examples include Applied Materials’ move from a Hadoop-based data lake to a governed lakehouse, with 1,500-plus analysts, more than 100 production machine-learning models and faster pipeline development, while Fonterra centralized data to improve supply-chain and compliance work. Rivian unified vehicle, factory and enterprise data, consolidated legacy systems into Delta and Unity Catalog, and designed for 500–600 petabytes; Lippert deployed AI tools across customer care, finance and HR. The announcement presents data and AI adoption as a cross-functional operating model.
Sara SteffenBring Databricks into Kiro IDE with the AI Dev Kit Power
AI-assisted development can produce unreliable SQL and models when an assistant guesses schema details or exceeds the user's data access. Kiro can connect to Databricks through Model Context Protocol (MCP) in two ways: four Databricks-managed remote servers for Genie, SQL, Unity Catalog Functions, and Vector Search, or the Databricks AI Dev Kit Power, which installs a local Python MCP server and broader skills. The AI Dev Kit now supports Kiro through its unified installer, while the Power provides one-click onboarding with authentication detection and skill loading. Both paths use Unity Catalog permissions, including row-, column-, and tag-based grants, so the assistant sees the user's effective access; Path A uses token-based configuration, while Path B supports OAuth U2M, OAuth M2M, profiles, or PATs. Examples show schema-grounded SQL, dbt joins using real columns, query comparisons, lineage checks, and generation of Databricks jobs or Asset Bundles.
Antony Prasad Thevaraj, Venkatavaradhan ViswanathanScaling Enterprise Conversational Intelligence: Cross-industry Technology and Functional Solutions Powered by Databricks Genie
Databricks Genie is presented as a cross-industry technology layer for enterprise challenges including financial planning, legal compliance, and IT operations. As a “Research Agent,” it can generate multi-step research plans to explain business anomalies and support answers with verifiable proof from the lakehouse. The post showcases partner solutions across technology, sales, marketing, HR, finance and procurement, supply chain, customer service, and IT operations, with examples spanning governed analytics, multi-agent orchestration, data observability, causal analysis, and incident management. These implementations aim to replace fragmented or static workflows with real-time, contextualized intelligence and production-grade agentic workflows, supporting anomaly investigation, root-cause analysis, ticket classification, and conversational troubleshooting. The stated goal is faster, more confident decision making across departments through governed self-service access to insights.
Amit SinghIntroducing Cross-Engine ABAC
Cross-engine ABAC is announced in Beta, extending Unity Catalog's fine-grained governance to external engines through Iceberg REST Catalog APIs. It supports tag-based row filters and column masks, including conditional logic and SQL UDFs, while allowing policies to be defined once and enforced across engines. For an external query, the engine sends a scan request, Unity Catalog evaluates entitlements and applicable policies, and returns a filtered scan plan before the engine processes authorized files. Enforcement remains at the catalog layer, so engines need not implement governance logic and can use the open scan APIs. Apache Spark is supported today through Iceberg-Spark and Delta-Spark connectors, with Starburst and DuckDB integrations coming soon; the Beta also points toward Apache Iceberg label exchange for future governance metadata sharing.
Alex Jiang, Alex Reid, Michelle LeonAdvancing Apache Iceberg on Databricks: Iceberg v3 GA, Open Sharing, and Unified Governance
Databricks announces a broad set of Apache Iceberg capabilities in Unity Catalog, spanning General Availability, previews, and beta releases. Managed Iceberg is GA, supporting table creation, reads, writes, optimization, governance, and sharing, while Iceberg v3 adds deletion vectors, row tracking, and VARIANT across managed, foreign, and UniForm-enabled tables. Unity Catalog also federates external catalogs, vends credentials, shares live data with Iceberg REST-compatible clients through Delta Sharing, and applies attribute-based access control during server-side scan planning for supported external engines. These capabilities are presented as a unified approach to open APIs, cross-engine governance, zero-copy sharing, and production performance without copying data. The post also outlines Iceberg v4 and a proposal for Delta 5.0 to adopt an adaptive metadata tree structure.
Jason Reid, Ryan Blue, Daniel Weeks, Michelle LeonBI Serving Pointers; Maximizing for Performance and TCO
BI dashboards can become slow and expensive when teams respond to latency with separate aggregate tables, refresh pipelines, extracts, and tool-specific semantic layers. The post presents Databricks’ BI serving stack from physical storage through Unity Catalog’s governed semantic layer, recommending Gold-layer star schemas, managed tables, liquid clustering, and Predictive Optimization to reduce scanned data and improve query plans. Metric Views centralize KPI definitions and semantic metadata for dashboards, Genie, SQL notebooks, third-party BI tools, and AI agents, while materialization automatically maintains incremental pre-aggregations and routes queries transparently. Additional TCO guidance covers serverless SQL warehouse autoscaling, DBSQL disk and query-result caching, direct lakehouse connections, and system-table monitoring. The stated outcome is compounded lower latency and compute cost, including an observed average 22% performance improvement from Predictive Optimization and sub-second performance from materialized metrics.
Chris KoesterAnnouncing Lakebase Change Data Feed (CDF)
Lakebase Change Data Feed (CDF) is available in Public Preview to reduce the manual effort of moving data from operational databases into downstream systems. The feed is enabled once for all tables in a project, stored and governed in Unity Catalog Managed Tables, and readable by engines, models, and agents without separate extraction pipelines. From one shared feed, teams can build streaming pipelines with SDP, create materialized views with DBSQL, or compute and store embeddings with Agent Bricks, while consumers remain isolated from the primary operational workload. The announcement positions Lakebase as the native Bronze layer in a medallion architecture, complementing Synced Tables and providing governance and lineage across the data lifecycle.
Pranav Aurora, Cheng Chen, Hristo StoyanovAI readiness in telecommunications
Telecommunications companies are adopting AI for customer experience, network operations, and cost reduction, yet initiatives often stall before production because fragmented, ungoverned, semantically opaque data creates data debt. The post argues that AI readiness depends on a semantic layer unifying datasets and business definitions, governance, and catalog metadata across systems such as Oracle, Snowflake, Salesforce, ServiceNow, and Databricks. It presents Unity Catalog as the proposed foundation, using Delta Sharing, Lakeflow Connectors, and Lakehouse Federation to exchange, ingest, or query data without uniformly replicating it, while privilege-aware metadata and audit logging support compliance. Metric Views, lineage, tags, and glossaries give agents authoritative meanings for measures and terms such as revenue, ARPU, active user, and FTTH. The conclusion is that trustworthy operational AI requires a governed, unified data foundation and organizational commitment, not simply more capable models.
Stephen Hage, Keerthi Josyula, Michael ZhangObservability for any agent, anywhere: Production-ready tracing with OpenTelemetry & Unity Catalog on Databricks
Databricks supports writing OpenTelemetry (OTel) traces directly to Unity Catalog, where real-time telemetry is stored in Delta tables for governed analytics and retention. AI traces capture prompts, tool calls, responses, latency, and execution paths, enabling debugging, evaluation, and monitoring, while lakehouse storage also allows SQL queries, dashboards, joins with business data, and PII controls. The managed, serverless ingestion layer uses Zerobus Ingest to accept OTLP over gRPC from collectors and REST integrations, streaming spans, logs, and metrics to Unity Catalog without intermediate message buses. A LangGraph support manager assistant demonstrates instrumentation with mlflow.langchain.autolog() and an @MLflow.trace root span, while Genie is invoked through MCP for data-driven questions. The resulting traces can be searched in MLflow, evaluated at scale, and monitored continuously, with FAQ details stating 200 QPS starting throughput, no storage limit, and Unity Catalog governance options for access control, masking, and filtering.
Firas Farah, Bruno Faria, Anoop SunkeHow Databricks Genie democratizes data access in financial services
Financial services organizations have built sophisticated lakehouses, streaming pipelines, model-serving infrastructure, and self-service BI, but access remains concentrated among technical teams. Business leaders still often rely on analysts because they may lack SQL skills, BI training, or analyst access, creating the “last mile” of data democratization. Databricks Genie addresses this gap through a conversational AI interface that converts plain-English questions into governed SQL queries executed against the Databricks Lakehouse without an analyst in the loop. It operates within Unity Catalog access policies, restricts users to authorized data, makes queries read-only, and logs interactions for audit purposes, while its semantic layer maps organizational terminology such as NIM, LTV, and NII to the organization’s meanings. The stated outcome is faster, auditable answers for business questions and usage data that can inform data-product priorities.
Kim HattonTransforming industries with conversational AI: Partner solutions built on Databricks Genie
Databricks presents a first cohort of partner-built industry solutions that use Genie to reduce the “Analyst Bottleneck,” where leaders wait for custom SQL or dashboard updates. Genie provides a conversational analytics layer grounded in Unity Catalog metadata and business semantics, using specialized AI agents to return governed, secure, traceable answers to natural-language questions. The featured accelerators span communications, media and entertainment, financial services, healthcare and life sciences, manufacturing and energy, public sector, and retail, with examples including churn prediction, next-best-action recommendations, billing anomaly investigation, KYC monitoring, actuarial analysis, and commercial insights. Across the examples, solutions combine domain expertise with Databricks platform capabilities such as machine learning, Agent Bricks, Lakebase, Serverless SQL, and auditable governance to support self-service analysis and operational decisions.
Amit SinghYou’ve built the media products, now make them personalized
Media companies may have launched streaming services, digital editions, and mobile apps, yet still face a digital product intelligence gap: product teams must answer behavioral data questions quickly enough to personalize experiences and improve engagement. Databricks Genie gives nontechnical product leaders a conversational interface that translates natural-language questions into SQL queries, visualizations, and actionable insights over governed enterprise data, without requiring code or analyst handoffs. The agent queries governed Delta Lake tables fed by streaming clickstreams, views, and session signals, combining event-level behavior, A/B tests, audience segments, and cross-platform data in conversational answers. The source says internal benchmark accuracy improved from 32% to over 90% through multi-LLM orchestration, specialized knowledge search, and parallel reasoning, and presents Genie as reliable enough for production personalization decisions.
Elena TesserFrom "What Happened?" to "What Will Happen?"
Databricks Genie makes descriptive analytics accessible in natural language, but predictive questions still require specialized data science workflows and carefully prepared datasets. This post presents a multi-agent supervisor deployed as a Databricks App, combining Genie, TabPFN, and Agent Bricks to turn business questions into predictions. The orchestrator asks Genie to use governed Lakehouse data, schemas, relationships, and semantics to generate labeled training data through SQL, then sends it to TabPFN, which predicts in a single forward pass without feature preprocessing, model selection, or hyperparameter tuning. The resulting conversational experience supports descriptive and predictive analytics with Unity Catalog lineage and access control, while an MLflow GenAI evaluation harness monitors reliability and regressions. Its central limitation is that predictions depend on Genie producing a meaningful dataset with a clear label, so missing signals, joins, outcomes, or agent omissions can make results unreliable.
Ryuta Yoshimatsu, Javier Poveda Panter, Dominik Safaric, Philipp Singer, Diana Kriuchkova, Sauraj Gambhir, Dael Williamson, Bryan Smith