Loading…

Databricks
Data and AI platform for data engineering, analytics, machine learning, and generative AI.
Latest articles
AI Transparency: Governance, Explainability, and Data Practices
AI transparency is presented as the practice of documenting and disclosing an AI system’s data, model behavior, decision-making processes, and accountability so affected stakeholders can evaluate its outputs. The guide distinguishes transparency from explainability, which addresses a specific prediction, and interpretability, which concerns direct access to a model’s internal logic. It recommends maintaining model architecture and version history, algorithms and hyperparameters, training-data provenance, and explainability tooling in a central registry, alongside model cards and data sheets. It also calls for subgroup performance metrics, recurring audits, visible AI disclosures, human-review paths, and incident playbooks covering notification, logs, rollback, remediation, and outcome tracking. The stated goal is durable governance that supports trust, bias detection, regulatory documentation, and accountability across high-stakes and generative-AI deployments.
Databricks StaffResponsible AI: Governance, Principles, and Practical Guide
Responsible AI is presented as a lifecycle-wide practice for designing, developing, deploying, and monitoring AI systems with fairness, transparency, accountability, privacy, safety, and human oversight as requirements. The guide connects technical controls—secure encrypted data pipelines, documented dataset provenance, demographic bias audits, adversarial robustness tests, access controls, and continuous monitoring—with governance mechanisms including named model owners, cross-functional oversight, model-risk assessments, and immutable decision logs. For generative AI, it recommends output policies, training-data leakage testing, guardrails, and red-team testing, while model cards, automated fairness checks, independent audits, and incident response plans support transparency and accountability. Regulatory preparation includes mapping systems to the EU AI Act’s risk categories and documenting design, training data, and intended use; the NIST AI Risk Management Framework and OECD AI Principles are identified as governance references.
Databricks StaffTech builds on AI. Finance protects the margin.
AI-native tech companies must protect unit economics as agents accelerate changes in compute consumption, pricing, and revenue recognition, while gross margins remain below classic software levels. Finance teams built around extracts, spreadsheets, and monthly reconciliation can miss repricing changes, metering errors, and compute-commitment risk. The proposed foundation is an evolving ontology that keeps product, plan, usage, and cost meanings current, with Stripe data entering Unity Catalog through OpenSharing and Lakebase providing transactional Postgres on the lakehouse. Genie One uses that ontology to answer governed, sourced questions about gross margin, consumption revenue at risk, and compute spend, while people retain decision authority. The post describes organizations using Databricks to consolidate reporting, forecasting, workflows, and finance applications, positioning a shared data-and-AI platform as the path from an initial answer to an ongoing finance platform.
Madelyn MullenBuilding a soccer coaching app on Databricks
Coach’s Corner, also called La Pizarra, turns high-frequency soccer tracking data into a bench-side application for replay, tactical analysis, scouting, standings, and agent-generated dossiers. Built as a Databricks App, it ingests NDJSON feeds at 25 frames per second through Auto Loader and Spark Declarative Pipelines, enforcing 46 data-quality expectations across bronze, silver, and gold layers. Liquid clustering supports 1–3-second DBSQL queries, while Lakebase synchronizes gold data to Postgres for millisecond replay reads and separates sequential playback from exploratory analytics. The scouting layer grounds Genie, Vector Search, a Unity Catalog-registered xG model, and an Agent Bricks supervisor in governed data, with Claude calls routed through the Unity AI Gateway, MLflow tracing, and a deterministic fallback. Together, these components are presented as a way to deliver traceable insights within seconds without forcing coaches to interpret raw tables or analysts to relay every result.
Samwel Emmanuel, Sheridan Harris, Andrew Helmreich, Kush Patel, Nick RagoneseMeta’s Spark Muse 1.1 is now available on Databricks, fully governed by Unity AI Gateway
Databricks announces support for Meta’s Muse Spark 1.1 through Model Provider Services (MPS) in Unity AI Gateway, addressing fragmented API keys, access controls, and usage visibility when organizations adopt new models. An MPS is a Unity Catalog securable that stores provider configuration and an encrypted API key, while callers use their Databricks credentials and the gateway attaches the key at request time. The post demonstrates registering Muse Spark through the OpenAI provider type with Meta’s API base URL and Responses API, then governing use with Unity Catalog privileges, model allowlists, policies, rate limits, usage metering, and inference tables. Requests are routed through the gateway, where access and guardrails are applied before reaching Meta; usage, spend, tokens, latency, status codes, and optionally full payloads are recorded for attribution and audit.
Pavithra Rao, Shaotong Li, Martin Grund, Kelly AlbanoYour AI is ready. Your data foundation probably isn’t
Cushman & Wakefield’s enterprise AI program addresses fragmented experiments, disconnected data, and uneven organizational maturity across a 53,000-person workforce. Over four years, Chief Digital and Information Officer Sal Companieh used a product operating model, business-linked accountability, co-created investment decisions, and shared architecture standards to build a common foundation while preserving business-unit flexibility. Databricks supports that strategy as a partner and platform, with its intelligence layer and Genie helping employees query trusted data in natural language, examine quality and governance, and monitor compliance. The company says the time from idea to outcome has fallen from months to days, while client and acquisition onboarding has materially accelerated. Companieh identifies human behavior, education, and trust—not technology alone—as essential to making change durable.
CIO.comFrom experiment to insight: how Dotmatics Luma and Databricks make AI-ready science a reality
Scientific workflows generate data across instruments, teams, and partner networks, but siloed systems can strip away metadata, lineage, and experimental context. Luma, Dotmatics’ scientific intelligence platform, captures outputs continuously and harmonizes them into structured, FAIR-compliant records, while Databricks supplies scalable storage, governance, and enterprise data and AI infrastructure. Running natively on Databricks, Luma preserves a continuous digital thread across experiment design, acquisition, analysis, reporting, and sharing, with Delta Sharing supporting governed exchange with collaborators. Chromatography illustrates the approach: connectivity across vendor systems and Virscidian’s Analytical Studio automate LC/MS processing while retaining context and adding dashboards, registration, and management tools. In one pharmaceutical deployment, Luma connected about 1,500 of more than 5,000 instruments across four LC/MS vendors, enabling cross-vendor performance trending, unified purity analysis, and a foundation for AI and machine learning.
Ryan Bernhardt, Michael FritzThe skills gap behind agentic AI — and how Databricks is closing it with a new context engineer certification and agent trainings
Databricks introduced the Databricks Certified Context Engineer Associate beta exam to validate skills for building reliable, production-grade AI agent systems as organizations scale agentic AI. Context engineering is presented as the practice of curating, maintaining, and filtering tokens, memory banks, and tool parameters so an LLM can solve a task; the certification is intended to benchmark that technical skill set. Databricks also added AI Agent Fundamentals, Building Retrieval Agents on Databricks, and Agent Evaluation, covering agent reasoning, retrieval-augmented architectures, and systematic performance testing and improvement. Its AI-first certification-prep guide is available on every certification page, works with free tiers of major LLMs, includes guardrails and disclaimers about hallucinations, and shows candidates how to use Free Edition for hands-on practice; registration for the Context Engineer Associate is open, with the first exam scheduled for July 29, 2026.
Rachel Canetta, Trang LeUnified context: The missing layer for enterprise AI coworkers
Enterprise AI assistants often produce fluent answers yet fail to improve forecast calls, deal reviews, and operational standups because decision context is scattered across systems, teams, and competing definitions. Genie One addresses this by using a shared context layer spanning Databricks data, documents, SaaS applications, and operational systems, allowing questions and follow-up work to retain business meaning. Genie Ontology organizes terms, metrics, entities, and relationships into a living knowledge graph, learning from data, dashboards, queries, documents, and connected applications while ranking definitions and signals using usage and certified-asset links. Together with Unity Catalog, it applies permissions, certified data, shared definitions, and governance controls to answers, actions, and agents. The stated outcome is faster movement from decision preparation to action, with less manual reconciliation while preserving accuracy and control.
Cynthya Peranandam, Christy MaverWhat happens in the milliseconds after you tap pay
At checkout, fraud scoring must combine transaction inference with customer-specific rules while keeping latency low enough for an interactive payment. The retail-app sample pairs a FastAPI backend and React frontend in a Databricks App with Model Serving route optimization and Lakebase Postgres, where the model retrieves historical features and the backend reads profile controls. A transaction is scored first, then checked against daily spending, international-transaction, and country rules; pooled OAuth-authenticated connections and token-rotation handling avoid repeated handshakes and stale credentials. In the supplied benchmark, route-optimized calls reached 27.2 ms p50 and 37.3 ms p95 end-to-end, with feature lookup at 8.9 ms p50, CatBoost inference at 0.4 ms, and network overhead at 17.4 ms p50; actual performance varies by model.
Harsha Pasala, Subhadip ChandaHow Unity Catalog managed tables bring interoperability, performance, and unified governance to the Lakehouse
Unity Catalog external access to Unity Catalog managed Delta tables is now in Public Preview, allowing external engines to create, read, and write while governance remains centralized. The announcement addresses the previous multi-engine trade-off: external tables enabled access but lacked managed-table performance optimizations and governance guarantees. Catalog commits coordinate writes through Unity Catalog, making it the source of truth for table state and enabling safe external writes, multi-statement transactions, and auditing of external operations. Predictive Optimization cleans storage, collects query statistics, and selects Liquid clustering columns as query patterns change; the post says these capabilities can deliver up to 50% storage cost savings and 20x faster queries. Support includes Spark, Flink, Starburst, DuckDB, and StreamNative, with open APIs and Delta Kernel extending integrations across Databricks UC and UC OSS.
Alex Jiang, Tathagata “TD” DasIntroducing Apache Spark 4.2
Apache Spark 4.2 extends the engine’s role in modern data and AI workloads with governed metrics, vector and top-K primitives, Arrow-first Python execution, native change data capture, and stronger streaming foundations. Metric views provide shared business definitions, while Spark Connect uses gRPC and Arrow to let remote clients submit logical plans without a full Spark runtime. Spark SQL adds vector similarity functions, NEAREST BY, geospatial types, sketches, and time-series features; Python interoperability includes Arrow UDFs and can move Spark DataFrames into supported Arrow-native tools without copying or serializing underlying data. Spark Declarative Pipelines adds Auto CDC for SCD Type 1 targets, while Data Source V2 standardizes change streams through CHANGES and expands row-level operations and schema evolution. The release also includes operational updates such as Web UI modernization, Kubernetes improvements, JDK 25 support, and dependency upgrades.
Wenchen Fan, Andreas Neumann, Serge Rielau, Szehon Ho, Gengliang Wang, Linhong Liu, Hyukjin Kwon, Jerry Peng, DB Tsai, Xiao Li, Reynold XinInkling model from Thinking Machines Lab now on Databricks
Databricks announces that Inkling, Thinking Machines Lab’s first open-weights model, is available to enterprise customers through the Unity AI Gateway. The model is positioned for coding and agentic reasoning workflows, supports multi-modal inputs, and can be applied to enterprise data, including proprietary codebases, internal documentation, and domain-specific data. Unity AI Gateway provides centralized security, permissions, audit logging, policy enforcement, cost controls, budgets, and observability, while data remains within the governed environment; Inkling is invoked through a REST API, with SQL query support planned. Teams can try it in AI Playground, deploy a governed endpoint, connect coding agents such as Cursor, OpenCode, or Pi, and build agents with Agent Bricks, with open weights enabling customization and inference-cost optimization without per-token API pricing.
Mike Eastham, Yuchen Jin, Preslav LeAI-Enabled Advisory Services for Higher Education
Higher-education call centers face costly, limited-coverage monitoring of advisor conversations and brittle, slow methods for identifying student concerns from transcripts. The proposed workflow deploys OpenAI Whisper on Databricks Model Serving, applies AI Functions for sentiment, topics, intent, and rubric scoring, and uses Unity Catalog to govern the resulting data. For advisor quality, an LLM-as-a-judge evaluates every transcript against a reference-table rubric, returns a weighted 1–5 overall score and per-criterion scores, and routes flagged calls for targeted QA review instead of random sampling. For student insights, quarterly transcript enrichment feeds an Agent Bricks Knowledge Assistant for cited reasoning over raw calls and a Genie Space for structured trend queries, while LangGraph orchestrates UC SQL functions as tools. Together, these components let non-technical advisors, mentors, and QA managers query student interactions without reaching out to a data SME.
Chad Ammirati, Zach Langford, Nicole WongData-Native AI Agents: Why Agents Must Move to Your Data
Enterprise AI pilots often move data into separate vector databases, SaaS LLMs, or serving layers, creating governance gaps, compounded latency, fragmented costs and observability, and duplicated lifecycle work. The post advocates data-native agents: models, agents, tools, retrieval, and memory run inside the governed data platform, with policy enforced during query planning and computation rather than after responses are produced. It argues that post-hoc controls cannot undo sensitive information encoded in aggregations and can trigger token-burning retry loops. For state and memory, it presents Lakebase, managed PostgreSQL within Databricks, as transactional storage and a shared source of truth for multi-agent swarms. The described platform pattern combines Unity Catalog, Unity AI Gateway, Model Serving, MLflow 3, AI Search, Lakebase, and business-context services, and recommends inventorying workloads already outside the perimeter before closing seams incrementally.
Kaan Kuguoglu, John KarlssonTake insights anywhere with Genie One on mobile
Genie One mobile apps for iOS and Android let business users ask questions of company data, view dashboards, access Databricks Apps, and use conversational agent capabilities away from a desk. Answers draw on Genie Ontology, Genie Agents, enterprise governance, and connectors to Google Drive, Microsoft 365, and Atlassian, while respecting source permissions. AI/BI Dashboards reflow widgets into a single column in portrait mode and preserve their designed layout in landscape. Authentication uses the existing OAuth flow and identity provider, with MFA, conditional access, device posture, network controls, regional data handling, and workspace permissions carried over from the browser; there is no separate mobile backend or mobile-only endpoint. The app is available in Public Preview, with dark mode, push notifications, voice mode, and account-level access listed as planned capabilities.
Mohit Hingorani, Christine Li, Rhetta Nadas, Sydney SundellHow Retail Finance teams are using Agentic AI to protect omni-channel margins
Omni-channel retail has spread margin, cash, and markdown decisions across more channels, fulfillment paths, and return routes, while agentic systems increase the speed and complexity of change. The post presents ontology as a way to preserve the meaning and business context behind finance figures, keeping definitions, channels, and cost drivers current. Databricks Genie is described as a data-smart AI coworker that answers finance questions in plain language, grounding responses in an evolving ontology, source traces, permissions, and governed AI costs. It focuses on margin after fulfillment and returns, inventory cash tied up in the wrong place, and full-price revenue at risk from markdowns and returns, then prepares actions for a person to approve. Unilever deployed Genie to more than 1,200 finance and business users; analysis that took days now takes minutes, with expected multi-million-euro annual cost avoidance.
Sarah DuffyFoundational context: Cross-industry & function-specific accelerators for Lakebase
Databricks presents Lakebase as a fully managed, serverless, standard Postgres database for combining operational and analytical workloads on its Data + AI Platform. The platform separates compute from storage, integrates with the lakehouse through Synced Tables and Lakebase CDF, and uses Unity Catalog for governance; copy-on-write branching and autoscaling to zero are described as core infrastructure primitives. The post showcases partner-built, ready-to-deploy accelerators spanning technology, finance, marketing, sales, supply chain, human resources, customer service, and operations. Examples include PostgreSQL migration assessment, multi-agent Genie orchestration, stateful enterprise agents, autonomous data reliability, governed contact-center intelligence, and project operations management. These offerings package Lakebase patterns into migration controls, domain-specific solutions, and agent frameworks intended to accelerate modernization and reduce transformation complexity.
Amit SinghBlocking Slow-Burn Attacks: Contextual Policies in Omnigent
Omnigent’s post examines how a vendor-review assistant can leak confidential pricing terms when an attacker hides an indirect prompt injection in a shared runbook. Because the malicious workflow is divided into ordinary actions, stateless checks approve each step even though the session as a whole is unsafe. The demonstration compares an unprotected run, which sends the summary externally, with a contextual policy that stores a running risk score, adds 30 for each document read, and denies email after the score exceeds 50. It also shows that agents cannot remove or disable policies, new policies require human approval, and any denial prevails when policies are combined. Runtime enforcement therefore preserves the block even when the agent has been misled.
Nishith Sinha, Matei ZahariaUltra-Fast Anomaly Detection using Apache Spark Real-Time Mode
The post presents a reusable real-time guardrail pattern for flagging suspicious Ethereum blockchain transactions and routing them for downstream action. Its rules identify impossible blocks where gas_used exceeds gas_limit and scan extra_data for email addresses, JWT tokens, or AWS access-key patterns, producing ALLOW or QUARANTINE decisions with reasons. The implementation uses Apache Spark Structured Streaming Real-Time Mode, whose continuous data flow, pipeline scheduling, streaming shuffle, pre-allocated execution pipelines, and asynchronous checkpointing target millisecond latency without a separate streaming engine. On a four-worker DBR 16.4 LTS cluster, the stateless Kafka-to-Kafka test processed about 23 million records at 69,713 rows per second, with P95 below 0.5 milliseconds and P99 at 1 millisecond. The post notes that RTM with a Kafka sink provides at-least-once delivery and that more complex stateful workloads may incur higher latency.
Jitesh Soni