Loading…
AI Agents
119 posts about AI Agents. Every summary links to the original.
AI governance at Data + AI Summit 2026: What’s new with Unity AI Gateway
Databricks announces new Unity AI Gateway capabilities for governing enterprise AI as organizations operate multi-model, multi-agent, and multi-vendor estates connected to models, MCP services, APIs, and tools. The update adds unified spend visibility, granular attribution, hard spend caps, and smart routing, alongside Unity Catalog support for registering and governing models, MCP services, agents, and skills. Contextual Service Policies, in Beta, can allow, deny, or require approval for actions based on users, agents, models, tools, services, or request and response contents, with guardrails for risks such as PII exposure and prompt injection. The announcement also covers end-to-end tracing, coding-agent analysis with Genie, incident investigation with Lakewatch, ecosystem integrations, and Managed Omnigent on Databricks in Beta.
David Nasi, Stefania Leone, Ahmed Bilal, Kevin Stumpf, Martin Grund, Vladimir Kolovski, Kelly AlbanoLakeflow: A new era of agentic data engineering
Databricks announces a major evolution of Lakeflow, its unified platform for data engineering across ingestion, transformation, and orchestration, with capabilities centrally governed by Unity Catalog. Genie Code and generally available Lakeflow Designer support agentic and no-code pipeline development, while Genie ZeroOps monitors production assets, analyzes failures, proposes fixes, and validates them in a governed sandbox before human approval. Lakeflow Connect expands to more than 100 managed connectors, and Zerobus Ingest adds Kafka-compatible, gRPC, REST, SDK, and OpenTelemetry interfaces for high-volume event ingestion. Real-Time Mode for Spark Declarative Pipelines reaches Public Preview with end-to-end latency as low as 5 milliseconds, alongside declarative APIs and expanded Lakeflow Jobs integrations. The release also adds data-readiness triggers and external orchestration for systems including Snowflake, REST APIs, Slack, and PagerDuty.
Bilal Aslam, Ray Zhu, Manish Dalwadi, Saad Ansari, Giselle GoicocheaIntroducing Genie ZeroOps: Put your data and AI operations on autopilot
Genie ZeroOps is an autonomous background agent for monitoring and operating data and AI assets, including jobs, pipelines, tables, and ML models. It continuously detects visible and silent failures, uses Unity Catalog lineage and platform observability to assess root causes, generates remediation through development workflows, and verifies fixes in isolated sandboxes. These environments use shallow, zero-copy table clones, scoped permissions, and network isolation, so proposed changes run against real data without touching production or applying anything before approval. For ML workloads, the agent can diagnose degraded predictions, train a candidate on corrected features, evaluate it against the production model’s existing eval suite and criteria, and support live-traffic ramping when it is measurably better. Genie ZeroOps is entering private preview in the coming weeks, initially supporting jobs, pipelines, tables, and ML workloads; Apps and Lakebase databases are on the roadmap.
Bilal Aslam, Lennart Kats, Ray Zhu, Mike Del Balso, Ori ZoharAnnouncing Lakebase Search: agent-native retrieval built into Lakebase Postgres
Lakebase Search is a beta offering on AWS and Azure that adds hybrid vector and full-text retrieval to Lakebase Postgres. It uses the lakebase_vector and lakebase_text extensions to keep retrieval, memory, operational data, and hybrid search in one backend. lakebase_vector retains pgvector types and operators, applies RaBitQ clustering and compression for 32x smaller indexes, and targets more than 1B vectors, while lakebase_text replaces GIN with object-storage-optimized BM25 ranking. A tiered cache keeps hot data on NVMe and places colder data in object storage; the source reports lower memory needs, faster index builds, and cold-cache startup than standard pgvector HNSW in its LAION-100M benchmark. The extensions also combine vector similarity and keyword relevance with reciprocal rank fusion in a single SQL query, enabling joins and tenant filtering alongside transactional workflows.
Pranav Aurora, Zhou Sun, Jinjing ZhouIntroducing Genie One, Genie Agents, and Genie Ontology
Databricks announces Genie One, Genie Agents, and Genie Ontology to help enterprises answer business questions and act on data whose context is scattered across dashboards, queries, documents, tickets, and chats. Genie One connects data and business tools through Lakehouse federation, Lakeflow Connect, native integrations, Slack, Teams, mobile apps, schedules, alerts, document creation, custom skills, and MCP support. Genie Agents evolve Genie Spaces into domain-specific agents that can reason over structured and unstructured data, execute multi-step workflows, and be created from a prompt. Genie Ontology builds a permission-aware living graph from enterprise assets, weighting sources by authority, usage, certification, and freshness. In an internal 28-question benchmark, Genie answered 84.5% correctly on the first attempt and delivered twice the speed of the strongest coding agent.
Sydney Sundell, Ken Wong, Elise GeorisIntroducing CustomerLake: The Agentic CDP embedded in Databricks
Databricks announces CustomerLake, an Agentic Customer Data Platform embedded natively in its lakehouse, bringing Customer 360, identity resolution, audience building, campaign automation, activation, and personalization alongside governed data and AI models. The announcement addresses fragmented identities, stale audiences, manual campaign workflows, and the duplication and governance burden created by separate martech systems. CustomerLake uses Unity Catalog and Lakehouse Federation to access customer data across Databricks, Snowflake, Google BigQuery, cloud object storage, operational databases, and other enterprise systems, while Profile Agents create business-ready profiles and Campaign Agents build audiences, recommend actions, activate channels, and optimize engagement. Its operating model is described as embedded, democratized, and autonomous, with Agentic Identity Resolution combining deterministic, probabilistic, and agentic workflows. CustomerLake is now available in Private Preview and launches with an open partner ecosystem.
Tasso Argyros, Justin DeBrabant, Michael Trapani, Dan Morris, Katy YuanIntroducing Omnigent: A Meta-Harness to Combine, Control and Share Your Agents
Omnigent is an open-source meta-harness designed to combine, control, and share agents across different harnesses, models, and interfaces. Databricks says users currently juggle multiple agents and copy context between them, while builders struggle to combine or replace harnesses with incompatible interfaces. Its runner wraps terminal-based agents and SDKs in sandboxed sessions with a uniform API for messages, files, streamed text, and tool calls; a server adds policies, sharing, and access through terminal, web, mobile, native Mac OS, and APIs. Features include live collaboration, hosted sandbox execution, contextual security and cost policies, OS isolation, and multi-harness authoring. Released in alpha under Apache 2.0, Omnigent is intended to provide a durable layer above changing agents and harnesses.
Matei Zaharia, Kasey Uhlenhuth, Corey ZumarEnabling Evolutionary Database Development: Database branching with Lakebase, the conclusion
The post concludes a series on how copy-on-write database branching in Databricks Lakebase changes team-scale evolutionary database development without changing its underlying methodology. For a team of fifty developers, long-running tier branches and ephemeral feature branches form a parent-linked promotion hierarchy, replacing separately provisioned environment instances and enabling promotion by merge, rollback by repoint, and computable schema divergence. Governance is declared once and inherited per branch, with policies intended to prevent transitions that contradict the parent chain; Unity Catalog captures metadata for attribution and audit. The DBA's role becomes platform engineering, while agents operate inside an executable SCM state machine with documented inputs, outputs, schema validation, and enforced gates. An optional TDD layer adds dedicated roles, acceptance-criterion scenarios, RED-GREEN-REFACTOR cycles, and artifact contracts, and the conclusion presents the resulting workflow as operational for human and agent practitioners.
Pramod Sadalage, Kevin HartmanUnlocking semantics for AI: How Mercedes-Benz Korea built trusted “Talk to Data” at scale
Mercedes-Benz Korea piloted a “Talk to Data” architecture that extends its Databricks analytics foundation with a governed semantic layer for enterprise AI, rather than treating the effort as a chatbot project. The design moves Power BI DAX KPI logic into Unity Catalog Business Semantics and Metric Views, keeping sources, joins, measures, dimensions, comments, and synonyms alongside governed Lakehouse data. Genie spaces use curated metric views for domain questions, while Agent Bricks composes persona-based agents, with Unity Catalog enforcing row- and column-level access. An automated DAX-to-Metric-View transpiler parses semantic models, maps tables, generates draft definitions, flags non-automatable measures, and reports conversion gaps. The documented playbook combines gold-layer curation, KPI validation, regression testing, Genie optimization, persona agents, and Databricks Apps; the pilot reports AI answers aligned with established KPI definitions and BI reporting logic.
Sai Yang, Fares Kamal, Alina Kamal, Andreas Jäck, Johannes Laufer, Manuel CulebrasEmpower your healthcare agents with ready-to-use MCP on Databricks Marketplace
Databricks announces ready-to-use Model Context Protocol (MCP) servers for healthcare and life sciences through Databricks Marketplace, addressing the need to combine curated biomedical knowledge, timely data, specialized tools, and private records. Listings include services for drug and target intelligence, literature, clinical trials, FDA information, Medicare coverage, ontologies, real-world evidence, clinical semantics, and interoperability, while Climb connects live public sources with private Gold-layer data under Unity Catalog governance. Marketplace and custom MCP servers are centralized in the MCP Catalog and governed by Unity AI Gateway, with Genie Spaces, AI Search, Unity Catalog functions, and SQL Warehouses available as managed MCP servers. Users can assemble agents in AI Playground, Agent Bricks, or notebooks, then deploy endpoints or apps with MLflow tracing, evaluation, human feedback, and AI guardrails; examples span molecular-property lookup, clinical questions, and drug research.
Yen Low, Mark Lee, Matthew Giglia, Nicholas Siebenlist, Jay Bhankharia, Paul Ford, Itai WeissStop building data products. Start building data services.
Howden’s rapid acquisition pace exposed limits in an enterprise data model built around one product per use case, downstream quality checks, and dashboard-driven consumption. Group Chief Data Officer Barry Panayi describes shifting to open, governed data services, moving mastering and quality checks closer to ingestion, and codifying reconciliation in the Accord data model. On Databricks, the company consolidated more than 100 sources of record, standardized pipelines and shared code, and built reusable assets for cross-domain analytics, while continuing to productionize models as consistent services. The account argues that AI agents require a composable services layer, and that insight lag—the time between data existing and being usable—matters more than freshness; conversational analytics through Genie also reduced dashboard-building work.
Aly McGueAWS and Databricks at Data + AI Summit 2026: Accelerating real-world AI innovation
AWS and Databricks describe their expanded collaboration at Data + AI Summit 2026, where AWS returns as a Legend Sponsor with sessions, demos, customer stories, and industry forums. The partnership centers on generative AI adoption, unified governance, and open data architectures, including an agentic stack that combines Amazon Bedrock, Bedrock AgentCore, Kiro, and the Databricks Data + AI Platform. A featured integration uses a governed MCP connection through Databricks Apps so AgentCore can query Unity Catalog-governed data, ask AI/BI Genie questions, and read low-latency state from Lakebase while honoring existing permissions. AWS will demonstrate these workflows at Booth #100 and present a session on federating Unity Catalog to AWS Glue, alongside customer examples including Mastercard, Talkdesk, nCino, Addepar, and Workday. Attendees can also join technical conversations, receptions, and a 14-day Databricks on AWS Marketplace trial with $400 in usage credits.
Sarah Jack, Taylor HossModern BSA/AML compliance on Databricks
AML operations are strained by fragmented systems, high false-positive volumes, manual case documentation, and opaque vendor scoring, leaving analysts focused on backlog rather than financial-crime intelligence. The proposed Databricks Data + AI Platform unifies transaction monitoring, KYC, sanctions, case history, and policy data under Unity Catalog, using Lakeflow Connect and a Bronze–Silver–Gold Delta architecture with masking, row-level security, and lineage. MLflow, Model Serving, Lakehouse Monitoring, and inference tables support institution-specific detection models, while Agent Bricks coordinates agents for evidence gathering, recommendations, and SAR drafting with analysts retaining final decisions. The architecture also uses Lakebase for governed operational state and Databricks Apps for analyst and executive experiences. Reported outcomes include a 75% reduction in false positives reaching the analyst queue and compressing three-to-six-hour investigations to minutes.
Kateryna Savchyn, Pavithra Rao, Mimi Park, Emerson BayukClaude Fable 5 is now available on Databricks, fully governed through Unity AI Gateway
Claude Fable 5 is now generally available on Databricks, with rollout across AWS, Azure, and Google Cloud through Unity AI Gateway. The Mythos-class model targets long-running, complex, and ambiguous work, including autonomous enterprise workflows, document question answering, code investigation, and multimodal tasks. In Databricks' OfficeQA Pro benchmark, Fable 5 achieved 57.9% correctness, setting a state of the art; compared with Claude Opus 4.8, it was 20% more accurate and used 12% fewer tool calls, but ran approximately 30% slower and generated 2.5x more output tokens. Unity AI Gateway provides unified API access, fine-grained permissions, Unity Catalog logging, request and tool-call guardrails, and spend controls. Agent Bricks supports domain-specific agents, while Anthropic's policy includes 30-day retention for trust and safety purposes only.
Ahmed Bilal, Ivan Zhou, Yash Oza, Gautam Venkatesh, Alice Li, Harish GaurTransforming solar and wind maintenance reports with Genie and AI agents
Plenitude and Databricks built an agent-based system that turns solar and wind plant maintenance PDFs into structured data for cross-plant analysis. Event-driven ingestion uses Databricks Jobs and the ai_parse_document AI Function to extract text, tables, figures, and metadata, then stores page- and object-level JSON records in Delta Lake with coordinates, version history, and links to source reports. A Genie space uses Unity Catalog metadata, knowledge-store instructions, and SQL generation to answer natural-language questions, produce visualizations, and export results, while Agent Bricks can orchestrate multi-step workflows and downstream actions. The design also applies automatic liquid clustering to dynamic queries and row-level security to restrict results by country. The resulting data layer supports historical trends, plant comparisons, recurring-fault analysis, and a foundation for predictive maintenance, although the source frames predictive use as a future improvement.
Maria VallarelliWhat is Human-in-the-Loop (HITL)?
Human-in-the-loop (HITL) is an AI and machine learning approach that places people in training, supervision, or decision-making to improve accuracy, safety, and ethical alignment. Its feedback loop can include data labeling, output review, escalation, approval, override, and continuous feedback, with confidence thresholds and risk scoring routing only selected decisions to people. The explainer distinguishes HITL, where review occurs before flagged actions, from human-on-the-loop monitoring and human-over-the-loop governance, and separates HITL from RLHF, a training-specific technique. It describes uses in medical imaging, moderation, autonomous vehicles, financial services, and AI agents handling consequential actions. Databricks Agent Bricks is presented as supporting governed traces and Agent Learning from Human Feedback, including a case where 32 feedback items improved instruction-following from roughly 12% to 80%.
Databricks StaffData + AI Summit 2026: Insider’s Guide for Financial Services Leaders
Data + AI Summit 2026 features a dedicated financial services program for leaders evaluating AI transformation across banking, payments, insurance, and professional services. It lists sessions on proprietary data for underwriting, responsible AI in banking and payments, and AI delivery in professional services featuring First American, American Modern Insurance Group, Vantage Bank Texas, Santander, FIS, Acxiom, EXL, Bain, and EY. The Financial Services Forum includes executive firesides with leaders from Morgan Stanley, JPMorganChase, Mastercard, and RBC Capital Markets. A financial services lounge at the Moscone Expo offers demos, Databricks experts, and Agentic Banker and Virtual CFO use cases. Training courses on AI Agents, Lakebase, and apps plus hands-on labs and certification sessions are presented as a route from strategy to execution with executives and technical teams splitting focus.
Kim HattonYour guide to the Telecommunications Industry Experience at Data and AI Summit 2026
Data + AI Summit 2026 presents a Telecommunications Industry Experience for operators responding to surging network traffic, regulatory pressure, cybersecurity threats, competition, and customer churn. The event, scheduled for June 15–18 in San Francisco, positions unified data and AI, governed workflows, and production use cases as the basis for operationalized, AI-native telecom models. Its June 17 Telecommunications Industry Forum features keynotes, presentations, and executive panels on customer experience, autonomous network operations, fraud prevention, secure agent deployment, and the return from modernizing legacy data warehouses. Breakout sessions cover automated metadata generation for Genie, conversational AI/BI, Lakeflow pipelines with Agent Bricks, and data exfiltration protection with egress monitoring. The industry lounge will demonstrate Model as a Service and agentic real-time decisioning, while the agenda emphasizes peer examples and architectural blueprints for scaling AI under telecom governance and compliance.
Elena Tesser, Nevash PillayRamp ·
Stack Benchmarking
Ramp’s Stack is an AI-native accounting suite designed to automate repetitive book-closing work, including reconciliations, variance analysis, data entry, and schedules and accruals. To avoid overfitting to individual design partners, Ramp built a benchmark of synthetic business worlds, realistic accounting tasks, and accountant-written grading criteria, including standard and roll-forward worlds for testing memory transfer. The benchmark contains 237 tasks and 3,469 grading criteria in the analyzed slice, and supports repeated runs to compare models, prompts, tools, skills, harness changes, and memory behavior. Optimization included ablating skills, shrinking a spreadsheet skill from 14,000 to 5,000 characters, and tuning the system against end-to-end task performance rather than narrow evaluations. The resulting Stack system achieved the highest agent performance, with 4% higher accuracy and 3% better Pass@1, while remaining on the latency frontier with GPT 5.4; schedules and accruals were harder for raw models than variance analysis.
Ryan StevensScaling Enterprise Conversational Intelligence: Cross-industry Technology and Functional Solutions Powered by Databricks Genie
Databricks Genie is presented as a cross-industry technology layer for enterprise challenges including financial planning, legal compliance, and IT operations. As a “Research Agent,” it can generate multi-step research plans to explain business anomalies and support answers with verifiable proof from the lakehouse. The post showcases partner solutions across technology, sales, marketing, HR, finance and procurement, supply chain, customer service, and IT operations, with examples spanning governed analytics, multi-agent orchestration, data observability, causal analysis, and incident management. These implementations aim to replace fragmented or static workflows with real-time, contextualized intelligence and production-grade agentic workflows, supporting anomaly investigation, root-cause analysis, ticket classification, and conversational troubleshooting. The stated goal is faster, more confident decision making across departments through governed self-service access to insights.
Amit Singh