Loading…

Databricks
Data and AI platform for data engineering, analytics, machine learning, and generative AI.
Latest articles
Ingest semi-structured data faster and more efficiently with Variant - Now Generally Available
Databricks announces general availability of the Variant data type for ingesting semi-structured JSON, XML, and CSV while avoiding the usual choice between flexible storage and fast, schematized queries. Variant Shredding, also generally available, stores common fields as columns in Parquet files, while Predictive Optimization uses workload and query patterns to identify important fields, collect statistics, and improve file skipping. More than 5,000 teams write Variant, with users executing over 500 million Variant queries monthly across more than 160 TB of data; shredding delivers nearly four-times-faster reads than unshredded Variant and 30-times-faster reads than JSON strings. Auto Loader and Zerobus can ingest Variant into Delta or Iceberg, and future plans include Liquid Clustering by Variant fields, expanded SQL functions, and additional integrations.
Jonathan Brito, Gene Pang, Harsh MotwaniDatabricks Completes Acquisition of Panther: Accelerating the Security Lakehouse Era
Databricks says it has completed its acquisition of Panther, an AI SOC platform, to accelerate its security lakehouse strategy. The combination pairs Lakewatch’s open, governed foundation for collecting, retaining, and analyzing petabyte-scale security telemetry with Panther’s operational SOC workflows and more than 100 out-of-the-box integrations. Panther adds detections-as-code, CI/CD-based authoring and deployment, and AI-native triage and investigation that correlate cloud, identity, SaaS, IT, and business data. The announcement presents the combined platform as a way to retain high-fidelity telemetry, preserve data ownership through open standards including OCSF, Spark, Unity Catalog, Delta, Parquet, and SQL, and automate investigations, detection refinement, and response workflows for modern security operations.
Andrew Krioukov, Jack Naglieri, Taylor Kain, Dave HerraldBackstage with Lakebase, part 3
This third part describes joining Backstage’s live operational Postgres catalog with analytical billing data so teams can query resource ownership and cloud spend together. Lakebase isolates compute per workload, allowing portal traffic and FinOps queries to use the same underlying storage; benchmarked catalog queries ran at 55–65 ms end-to-end and searches at two to four ms. Because Lakehouse Federation’s Postgres connector accepts static credentials while Lakebase applications use OAuth JWTs, the proof-of-concept creates a separate SCRAM-SHA-256 Postgres role for federation. This enables a single query to join Backstage resource names with Unity Catalog’s system.billing.usage without data movement, while branch billing remains visible per branch and endpoint. The operational considerations include password rotation, read-only grants, branch TTLs, and scale-to-zero endpoints that stop billing when idle.
Cameron Casher, Shanil Anushka Fernando, Kevin HartmanFoundations for an AI-forward healthcare organization
Healthcare organizations adopting AI face fragmented data, mismatched governance, and no repeatable operating model, rather than a shortage of ideas or vendors. The piece defines an AI-forward organization as one where AI can be built, trusted, and scaled through a foundation of unified data, visible guardrails, and an operating model that helps teams prioritize and move pilots into production. It describes patient identifiers differing across source systems, requiring manual reconciliation and creating recurring integration costs. Governance must avoid both untrusted outputs and approval processes so rigid that nothing leaves the sandbox, while self-service users need controlled access to clinical, operational, and financial data. The source says modern tooling can centralize permissions and enable a governed first-use case in days rather than quarters when scope and data are ready.
Ramiz Bozai, Sailesh Kadam, Kriti Sen Sharma, Andrew Wallace-Jackson, Grace CrispAgentic media buying cannot scale without the right foundation. See how buyers and sellers get there on Databricks.
Media buying remains slowed by fragmented coordination across emails, spreadsheets, PDFs, and phone calls, leaving teams to research inventory, compare pricing, negotiate, and issue orders manually. The post presents an agentic workflow in which buyer and seller agents discover inventory, negotiate prices, and book deals through IAB Tech Lab’s AAMP standards, including AdCOM, OpenDirect, OpenRTB deals, and registry-based discovery and trust. Its Databricks implementation runs self-contained applications on Databricks Apps, with CrewAI agent crews using Databricks Foundation Model APIs and Claude models, Lakebase providing transactional Postgres state, and MLflow tracing capturing decisions and tool calls. The example processes a $200,000 Q3 Brand Launch across CTV and Linear TV with a $38 CPM ceiling, and is packaged as a Databricks Automation Bundle accelerator deployable with a single command.
Joe Hu, Mandy Baker, Luke BarnesConvert proprietary code to open ANSI SQL with Genie Code
Databricks introduces an agentic converter in Genie Code to translate proprietary warehouse dialects into open ANSI SQL, initially supporting T-SQL, Snowflake, Redshift, Oracle, BigQuery, and Teradata. Migration projects provide a workspace hub for source files, conversion status, complexity scoring, dependency lineage, and collaboration, helping teams prioritize work and identify scripts that can move independently. When launched, swarms of subagents convert files in parallel, iteratively fixing errors and validating syntax and semantic intent; in the proof-of-concept, six of eight files converted successfully. Two files required review because stored procedures needed three-part Unity Catalog names, which could be fixed manually or encoded as a custom Genie Code skill for reuse across the codebase. The post also notes that Databricks supports multi-statement transactions, temporary tables, and stored procedures, while planned additions include legacy ETL sources, new target dialects, data migration, and reconciliation.
Jonathan BritoNBCUniversal’s Seamless Migration: Unlocking Scalable Analytics with Databricks
NBCUniversal sought to modernize its data infrastructure as growing data volumes, evolving analytics demands, and slot-based compute reservations created cost, contention, and flexibility concerns. Working with EXL and Databricks, it migrated workloads to a Databricks lakehouse architecture that unifies data engineering, analytics, machine learning, and business intelligence. The phased, risk-based program assessed workloads and dependencies, designed workspaces and governance with Unity Catalog, and used four custom accelerators for SQL translation, DAG migration, data transfer, and validation. Job-specific compute, autoscaling, Delta Lake layouts, MLflow, and Lakeflow Jobs supported workload-specific execution and operations. The reported outcome was a 30% reduction in data infrastructure costs, elastic scaling for high-traffic events, unified ML operations, and the onboarding of approximately 300 analysts to Databricks SQL.
Kevin Hill, Ludwig Kuznia, Anil Joshi, Jay Mehta, Manojit DanQuality care is the mission. Finance protects the margin.
Healthcare finance teams must protect margins amid rising medical costs, denied or underpaid claims, complex payment arrangements, and cash trapped in receivables, while fragmented systems and delayed information increase decision risk. The post presents ontology as a way to preserve the business meaning behind figures, linking numbers to service lines, payers, contracts, and changing context rather than treating accuracy alone as correctness. Databricks Genie is described as a governed, data-smart AI coworker that answers finance questions with sourced figures, traces them to their origins, respects permissions, and keeps a person in the loop. It applies this model to care costs exceeding reimbursement, revenue lost to denials and underpayments, and aging or unbilled receivables, with the intended actions of renegotiating rates, preventing claim errors, accelerating collections, and protecting margin.
Aaron ZavoraManufacturing runs on capital. Finance protects the margin.
Manufacturing finance must protect margin by keeping capital moving through inventory, receivables, and plant equipment. The post argues that volatile supply chains, shifting demand, rising costs, and increasingly influential agents make that work faster and more complex, while an estimated $1.7T remains trapped in excess working capital across large US companies. It presents ontology as a way to preserve the meaning and context behind figures, including plants, SKUs, customer terms, and changing business conditions. Databricks Genie is described as a data-smart AI coworker whose ontology learns from business systems and questions, produces sourced and governed answers, and helps finance identify trapped inventory cash, aging receivables, and underperforming assets. It readies actions such as releasing inventory, accelerating collections, or redeploying capital, while a person remains responsible for the decision.
Caitlin GordonEnergy runs on volatile markets. Finance protects the margin.
Energy finance must protect margin as hourly power and fuel prices, shifting forward curves, contract settlements, hedges, and major grid and generation investments make business conditions more complex. The post argues that accurate figures are not necessarily correct unless they retain context about the underlying asset, market, and contract, motivating a live ontology that keeps business meaning current. It presents Databricks Genie as a data-smart AI coworker that learns from enterprise systems and questions, returns sourced answers, preserves permissions, governs AI cost, and shows its work. The proposed uses cover live margin, revenue risk across PPAs and hedges, and funding AI-driven buildout, with a person making trading, dispatch, hedging, or capital decisions. Genie is described as helping finance identify risks early and prepare actions while keeping human owners accountable.
Caitlin GordonBringing real-time fraud prevention to government benefits
Federal benefits programs face an estimated $233 billion to $521 billion in annual fraud losses and roughly $186 billion in improper payments in fiscal 2025, while pre-payment vetting can delay legitimate aid. The post presents real-time detection as an alternative to “pay and chase,” combining rules engines, machine learning, adaptive and generative AI, entity resolution, behavioral analysis, and graph analytics to score claims before disbursement. It says Databricks can unify these layers and enable cross-agency sharing through OpenSharing, data masking, and Clean Rooms while retaining governance and data ownership. When fraud is detected, the platform can hold payments, add confirmed fraudsters to Do Not Pay, and assemble statute-linked evidence packets with digitally signed audit trails, while designated personnel authorize referrals. The source concludes that this model is ready for implementation at national scale, citing Databricks’ stated processing of 160 billion card payments annually in milliseconds.
Mike McWhorter, Johnathan Tafoya, Stephen HutsonAgents for production lines: Trusted decisions in real time
ProdLine CoPilot addresses in-shift production-line decisions when equipment faults threaten output and data is split across PLCs, SCADA, MES, ERP, and LIMS. The proposed system streams operational technology data into Databricks, joins it with slower business and quality records, and routes questions to domain specialists that calculate recovery, depletion, quality-risk, and maintenance signals. Its orchestrator reads current Unity Catalog state, uses SQL, Genie, AI Search, Model Serving, and Python-based tools, while MILP, stochastic scheduling, Monte Carlo, Bayesian risk, and Pareto methods provide analytical outputs. Recommendations are tested against 1,000 scheduling scenarios and returned with trade-offs in cost, overtime, throughput, quality, and service, but execution remains with line managers, quality leads, and maintenance. The demo drafts work orders, quality deviations, and schedule updates for approval, with traceability storing inputs, assumptions, constraints, approvers, and outcomes.
Mohammad KhelghatiHow agentic AI can help telecom finance teams protect the margin when every moment matters
Telecom finance teams face revenue leakage when services go unbilled or uncollected, fraud and partner-settlement errors drain charges, and customer churn removes future spending. The source argues that fragmented ordering, billing, ERP, and spreadsheet data makes month-end reconciliation too slow, turning discrepancies found weeks later into write-offs rather than recoverable revenue. It presents Databricks Genie as a data-smart AI coworker whose ontology captures business meaning across systems, stays current, and grounds sourced answers in governed data. Genie can answer billing, invoice, spend, and variance questions, surface anomalies, and support action while a person remains responsible for decisions; Lumen Technologies is cited as using it across the Office of the CFO. The text says Genie does not make billing, fraud, or pricing decisions.
Elena TesserHow NorthStar Anesthesia built a scheduling app for a workforce of 3,000 clinicians in weeks
NorthStar Anesthesia needed a better way for roughly 3,000 physicians and Certified Registered Nurse Anesthetists (CRNAs) to coordinate shifts, because its commercial scheduling platform hid colleagues’ time-off data and its dashboard pilot was not mobile-friendly. Synaptiq built a React TypeScript scheduling app with one software engineer and deployed it through Databricks Apps in a few weeks, using the existing Databricks data foundation, governance, and Microsoft Entra ID single sign-on. The app provides color-coded views by clinician type, facility selection, day/week/month calendars, shift notes, search and filtering, and time-off visibility, with data refreshing every 30 minutes. Daily unique users rose from 75–80 at launch to more than 110 as the first clinician group migrated, while users reported less stress; planned additions include natural-language queries, push notifications, and automated staffing reports.
Evan PandyaThe audience is the asset. Media finance teams need to understand them to protect the margin.
Media finance teams are tasked with understanding how audience value flows across subscriptions, advertising, and content investment so the business can protect margin. Streaming has turned one wholesale audience into multiple monetization strategies, while ad-supported tiers now account for 59% of new streaming sign-ups, making timely measurement and pricing analysis more important. Ontology preserves the meaning and context of figures as audiences, titles, channels, and business conditions change, distinguishing a correct answer from one that is merely accurate. Databricks Genie is described as a governed, data-smart AI coworker that answers sourced natural-language questions, learns from interactions, and shows its work; Genie-powered apps let hundreds of DIRECTV analysts and leaders query more than 1,200 customer-level attributes, while people retain decision authority.
Elena TesserChief data officer: role, responsibilities, and career guide
The chief data officer (CDO) is a senior executive who treats enterprise data as a strategic asset, overseeing data strategy, management, quality, governance, and its use in analytics and AI. The role sits between business strategy, information technology, and data science, while distinguishing the CDO’s focus on data from the CIO’s responsibility for technology infrastructure. Responsibilities span the data lifecycle, governance frameworks, access controls, quality metrics, business intelligence, analytics prioritization, and organizational data capability. The role has shifted from compliance-focused stewardship toward enterprise AI strategy, with enablement-oriented governance making reliable, AI-ready data accessible while protecting its integrity. The guide also presents executive authority, data literacy, and leadership judgment as important conditions for delivering measurable business value.
Databricks StaffFrom prototype to production: High QPS for Databricks AI Search
Databricks AI Search now offers generally available high-QPS scaling for Standard endpoints, addressing production workloads such as search bars, recommendations, and real-time entity resolution. Operators set a human-readable target_qps value when creating or updating an endpoint through the Python SDK, REST API, or UI, while Databricks provisions the required compute without manual replica counts, node sizing, or load balancers. Existing Unity Catalog governance and Delta Sync remain in place, and scaling progress is exposed through scaling_info as it moves from SCALING_CHANGE_IN_PROGRESS to SCALING_CHANGE_APPLIED; endpoint observability shows requests per second, latency, and health. Service principal authentication is recommended for high-QPS traffic, while personal access tokens are capped at a few tens of QPS, and new capacity applies when an index is created or synced.
Adam Gurary, Sheng Zhan, Ankit Vij, Vadim Antonov, Yu-Ju Huang, Dima KotlyarovHow Databricks manages its own coding agent spend with Unity AI Gateway Budgets
Databricks describes how it manages coding-agent costs as thousands of engineers use Claude Code, Codex, Cursor, and other tools, creating exposure to runaway automation and growing R&D spend. The company routes all agent traffic through Unity AI Gateway Budgets and separates short-term runaway-spend protection from long-term monthly spend governance. A daily budget triggers self-service acknowledgement through Slack, an internal portal, or the CLI, while a high monthly limit uses manager-approved, project-scoped tiers that expire. Both budgets apply simultaneously, so effective usage is capped by the lower of the month-to-date total plus one runaway increment and the monthly maximum. Centralizing metering also gives managers and finance shared usage data, and the company reports that approval queues disappeared, monthly requests became rare, and engineers stopped rationing usage.
Rohit Agrawal, Shuyu Cao, Darming Zhao, Zack Siegel, Aaron DavidsonGet Started with Genie One: Top AI Cowork Use Cases for Business Users
Genie One is presented as a data-smart, agentic coworker for business users who need work completed across calendars, CRMs, data warehouses, ticketing systems, and shared documents rather than through a standalone chatbot. The guide explains four workflows: recurring business reviews, meeting preparation and follow-up, knowledge-work document automation, and operational monitoring with alerts. Users connect relevant sources, provide templates or natural-language rules, create reusable skills, test outputs, schedule routines, and refine instructions over time; the monitoring workflow can compare metrics with defined thresholds and email alerts. The post recommends starting with one narrow recurring task, connecting its required sources securely with data or IT partners, and iterating toward a library of saved skills that can be used across the organization.
Cynthya Peranandam, Sydney SundellThe EU Digital Product Passport: a traceability deadline
The EU’s Ecodesign for Sustainable Products Regulation (ESPR) will require in-scope products sold in the single market to carry a machine-readable Digital Product Passport (DPP), with batteries first from February 2027. The passport links a unique identifier and data carrier to lifecycle information that must remain accurate, while operators retain responsibility for data stored on their own backend, tiered access, supplier inputs, and any Registry registration. Implementation is primarily a data-integration and governance problem, spanning tier-N suppliers, per-unit operational records, lineage, analytics, AI, and controlled sharing rather than QR-code generation. Databricks components including Lakebase, Unity Catalog, Lakeflow Spark Declarative Pipelines, Databricks Apps, and OpenSharing are presented with a battery-focused Solution Accelerator as an operator-side reference architecture. The stated benefits extend beyond market access to faster recalls, sourcing-risk visibility, and lineage-based sustainability reporting, but technology alone does not establish compliance.
Daniel Dahlin