---
title: "Data Governance"
description: "42 posts about Data Governance, summarised, each linking to the original."
---

# Data Governance
> 42 posts about Data Governance, summarised, each linking to the original.

## Articles

### [How the FDA is building a secure, AI-ready data foundation on Databricks for Government](https://yomu.fyi/post/how-the-fda-is-building-a-secure-ai-ready-data-foundation-on-databrick.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Filippo Seracini, Vijay Raja
- Published: Sep 1, 2026

The FDA built HALO (Harmonized AI and Lifecycle Operations for Data) as a secure, governed, AI-ready enterprise data platform for modernizing siloed systems without interrupting regulatory work. Its move to Databricks on AWS GovCloud, following FedRAMP High authorization sponsorship, added Unity Catalog as a governance layer across a multi-tenant architecture, while Terraform-based security patterns, PrivateLink, customer-managed keys, and the compliance security profile support regulated workloads. The agency migrated more than 5,000 users and 8,000 jobs and pipelines with zero downtime, refactoring over 1,000 pipelines and 4,000 notebooks. After onboarding eight centers and 30 programs, FDA reported query responses improving over 30%, compute costs falling over 20%, and provisioning and sharing time dropping over 75%. HALO also supports responsible AI use cases such as MARS, with humans retaining decision authority.


### [Operationalizing Genie Ontology in Your Data Stack](https://yomu.fyi/post/operationalizing-genie-ontology-in-your-data-stack.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Srujan Alase, Richard Tomlinson
- Published: Sep 1, 2026

Genie Ontology is presented as a way to give enterprise AI agents shared business context beyond a semantic model, including definitions, relationships, business rules, authoritative sources, and permissions. It combines Unity Catalog Semantics—Metric Views, Pages, and Domains—with context inferred from governed tables, queries, dashboards, notebooks, and other supported assets. The guidance recommends six progressive layers, beginning with clean gold data and resolved golden records, then metadata, semantic modeling, enterprise context, governance, and evaluation. Critical implementation details include declaring informational primary and foreign keys, defining canonical measures in Metric Views, adding synonyms and example queries, and using permissions plus human-reviewed automation. Rather than waiting for complete coverage, it advises starting with one high-value domain and metric, then using feedback, telemetry, benchmarks, and drift reviews to strengthen trust over time.


### [Data Governance Architecture: A Complete Blueprint for Modern Organizations](https://yomu.fyi/post/data-governance-architecture-a-complete-blueprint-for-modern-organizat.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 2, 2026

Data governance architecture is presented as a blueprint for aligning policies, roles, processes, and technologies with business outcomes. It defines objectives including consistent data definitions, data integrity, layered access controls, and secure self-service analytics, while assigning responsibilities across executives, architects, engineers, analysts, managers, and compliance teams. Core principles are accountability, transparency, consistency, and stewardship, supported by federated ownership through councils, data owners, and embedded stewards. The discussion compares DAMA-DMBOK, TOGAF, and Zachman according to organizational scale, regulatory context, and architecture maturity, and describes modern patterns including lakehouse, data mesh, and data fabric. It concludes that effective programs require executive sponsorship, documented roles, measurable quality controls, iterative implementation, and sustained change management.


### [Practical Data Warehouse Design and Architecture Guide](https://yomu.fyi/post/practical-data-warehouse-design-and-architecture-guide.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Databricks Staff
- Published: Jun 2, 2026

The guide presents data warehouse design as a business-aligned process for building, modernizing, or scaling analytics infrastructure, rather than merely storing data. It recommends defining analytics use cases and stakeholder needs first, then organizing a three-tier architecture of source, storage, and semantic output layers; cloud designs can decouple compute and storage and use open formats. A Bronze–Silver–Gold medallion flow preserves raw lineage, applies cleansing and deduplication, and produces consumption-ready dimensional models, while retention and archival policies control sprawl. For modeling, it favors star schemas for user-facing BI, uses snowflake normalization when redundancy is material, and stresses explicit fact-table granularity, domain-owned data marts, and workload-specific refresh cadences. Governance and operations include Unity Catalog, access controls, masking, lineage, multi-region deployment, disaster recovery, and CI/CD, followed by phased rollout through high-value domains.


### [Advancing Apache Iceberg on Databricks: Iceberg v3 GA, Open Sharing, and Unified Governance](https://yomu.fyi/post/advancing-apache-iceberg-on-databricks-iceberg-v3-ga-open-sharing-and.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Jason Reid, Ryan Blue, Daniel Weeks, Michelle Leon
- Published: May 28, 2026

Databricks announces a broad set of Apache Iceberg capabilities in Unity Catalog, spanning General Availability, previews, and beta releases. Managed Iceberg is GA, supporting table creation, reads, writes, optimization, governance, and sharing, while Iceberg v3 adds deletion vectors, row tracking, and VARIANT across managed, foreign, and UniForm-enabled tables. Unity Catalog also federates external catalogs, vends credentials, shares live data with Iceberg REST-compatible clients through Delta Sharing, and applies attribute-based access control during server-side scan planning for supported external engines. These capabilities are presented as a unified approach to open APIs, cross-engine governance, zero-copy sharing, and production performance without copying data. The post also outlines Iceberg v4 and a proposal for Delta 5.0 to adopt an adaptive metadata tree structure.


### [Building a FHIR-native health data platform on Databricks Lakebase](https://yomu.fyi/post/building-a-fhir-native-health-data-platform-on-databricks-lakebase.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Marcin Jimenez, Aleksandr Kislitsyn, Nikolai Ryzhikov
- Published: May 27, 2026

Healthcare organizations often keep a FHIR server, analytics warehouse, and ETL pipelines separate, replicating clinical data and fragmenting governance. The proposed architecture standardizes incoming HL7v2, C-CDA, X12, and proprietary data into FHIR using Health Samurai converters, terminology normalization, MDM/MPI deduplication, and Implementation Guide validation. Aidbox, Health Samurai's FHIR server and database, runs on Databricks Lakebase, while Moonlink synchronizes operational and analytical formats without ETL. This exposes one governed dataset through Spark, SQL, ML, AI/BI, FHIR API, SMART on FHIR, and SQL on FHIR ViewDefinitions. The architecture supports use cases including EHR optimization, value-based care, member engagement, and compliance capabilities, with insights connected to clinician and billing workflows through SMART on FHIR and CDS Hooks.


### [AI readiness in telecommunications](https://yomu.fyi/post/ai-readiness-in-telecommunications.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Stephen Hage, Keerthi Josyula, Michael Zhang
- Published: May 26, 2026

Telecommunications companies are adopting AI for customer experience, network operations, and cost reduction, yet initiatives often stall before production because fragmented, ungoverned, semantically opaque data creates data debt. The post argues that AI readiness depends on a semantic layer unifying datasets and business definitions, governance, and catalog metadata across systems such as Oracle, Snowflake, Salesforce, ServiceNow, and Databricks. It presents Unity Catalog as the proposed foundation, using Delta Sharing, Lakeflow Connectors, and Lakehouse Federation to exchange, ingest, or query data without uniformly replicating it, while privilege-aware metadata and audit logging support compliance. Metric Views, lineage, tags, and glossaries give agents authoritative meanings for measures and terms such as revenue, ARPU, active user, and FTTH. The conclusion is that trustworthy operational AI requires a governed, unified data foundation and organizational commitment, not simply more capable models.


### [Pharma launch analytics: How to compress the first 90 days and win the three years that follow](https://yomu.fyi/post/pharma-launch-analytics-how-to-compress-the-first-90-days-and-win-the.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Adam Crown
- Published: May 23, 2026

Pharmaceutical launch analytics depends on compressing the time between data signals and commercial decisions, because early choices shape a trajectory measured over 12-to-36 months. The source frames the first 90 days as three phases: weeks 1–4 validate feeds, set NBRx and patient-start benchmarks, and identify coverage gaps; weeks 5–8 support tactical adjustments through AI-generated narratives, adoption cohorts, and access-barrier escalation; weeks 9–12 recalibrate against benchmarks, shift promotional spend, and record decisions. Databricks Genie lets commercial leaders question unified Rx, specialty-pharmacy, payer-coverage, field-activity, and patient-services data in natural language at prescriber, territory, and regional granularity, with governance and benchmark context. The stated operating benefit is a decision cycle under seven days, enabling teams to detect suppression early, reallocate resources, and respond to access barriers while the launch remains correctable.


### [How Databricks Genie democratizes data access in financial services](https://yomu.fyi/post/how-databricks-genie-democratizes-data-access-in-financial-services.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Kim Hatton
- Published: May 22, 2026

Financial services organizations have built sophisticated lakehouses, streaming pipelines, model-serving infrastructure, and self-service BI, but access remains concentrated among technical teams. Business leaders still often rely on analysts because they may lack SQL skills, BI training, or analyst access, creating the “last mile” of data democratization. Databricks Genie addresses this gap through a conversational AI interface that converts plain-English questions into governed SQL queries executed against the Databricks Lakehouse without an analyst in the loop. It operates within Unity Catalog access policies, restricts users to authorized data, makes queries read-only, and logs interactions for audit purposes, while its semantic layer maps organizational terminology such as NIM, LTV, and NII to the organization’s meanings. The stated outcome is faster, auditable answers for business questions and usage data that can inform data-product priorities.


### [How security teams can report cyber risk to boards](https://yomu.fyi/post/how-security-teams-can-report-cyber-risk-to-boards.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Taylor Kain
- Published: May 22, 2026

Boards are seeking visibility into cyber risk, but technical reports often fail to connect security posture with business impact or financial exposure. The post explains that compliance and cyber risk leaders can use Databricks Genie to query vulnerability posture, asset criticality, threat intelligence, control data, and historical incident costs in a governed environment. It recommends probabilistic financial modeling, including Monte Carlo simulation, to run randomized attack scenarios and produce loss distributions; Value-at-Risk framing can make those results familiar to directors. This approach replaces qualitative red/amber/green reporting with expected-loss ranges, supports investment prioritization, and enables trend analysis and board-ready answers, while the suggested cadence combines quarterly strategic briefings, monthly operational reviews, and incident-triggered updates.


### [Expanded interoperability with Unity Catalog Open APIs](https://yomu.fyi/post/expanded-interoperability-with-unity-catalog-open-apis.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Alex Jiang, Tathagata Das
- Published: May 14, 2026

Unity Catalog’s expanded Open APIs address data silos by letting organizations use multiple compute engines while retaining centralized governance and a single copy of data. In beta, Apache Spark, Apache Flink, and DuckDB can create, read, write, and stream to or from UC managed Delta tables, with catalog commits providing serialized commits, transactional safety, and auditability. Delta Kernel, an open source Java and Rust library, abstracts low-level protocol details, helping connectors integrate external writes with catalog-managed commits while Predictive Optimization continues to run on accessed tables. Credential vending, now GA for tables, issues short-lived, scoped cloud credentials and supports M2M OAuth plus automatic refresh; volume credential vending is in Public Preview for unstructured data. The roadmap includes functionality for fine-grained row- and column-level ABAC on external reads, while external managed-table access remains in beta.


### [From manual to autonomous: how AI agents are transforming electric grid operations](https://yomu.fyi/post/from-manual-to-autonomous-how-ai-agents-are-transforming-electric-grid.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Julien Debard, Edward Tavares
- Published: May 14, 2026

Electric utilities are facing rising demand, retiring generation, extreme weather, aging infrastructure, and fragmented operational data that manual processes cannot manage at scale. AI agents are presented as a human-centered alternative that synthesizes heterogeneous data, learns from outcomes, and progresses from human-approved recommendations to exception-based control and eventually autonomous operations within defined parameters. Hawaiian Electric used a Retrieval Augmented Generation proof-of-concept with Databricks AI Search, Unity Catalog, and Lakeflow Declarative Pipelines to query regulatory documents and provide page-specific citations. The system reduced response times from five minutes to five seconds and was implemented in two weeks, while the article describes broader potential for predictive maintenance, outage response, load forecasting, and customer service.


### [Data quality is the AI strategy](https://yomu.fyi/post/data-quality-is-the-ai-strategy.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Aly McGue
- Published: May 13, 2026

NYU Langone Health’s AI strategy starts with data quality, arguing that healthcare AI cannot be reliable when source data is fragmented or inconsistent. The institution standardized on common transactional platforms, including one electronic health record and one ERP system, established authoritative data sources, and fixes data at the source rather than mapping it in the warehouse layer. Its Databricks-based unified data and AI platform, with Unity Catalog, supports clinicians, analysts, scientists, and corporate users across care, operations, and research, while real-time feeds power emergency-room decision-support models. Mherabi also describes a three-layer analytics model: structured visualizations, conversational tools such as Genie, and answers delivered in formats suited to the user. The stated conclusion is that upstream data discipline, governance, literacy, and adaptable platforms provide the foundation for trustworthy AI and timely clinical insight.


### [ABAC row filtering and column masking policies, governed tags, and data classification are now generally available in Unity Catalog](https://yomu.fyi/post/abac-row-filtering-and-column-masking-policies-governed-tags-and-data.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Adriana Ispas, Kristen Wilder, Jacqueline Li, Corey Sunwold, Menglei Sun, Viswesh Periyasamy
- Published: May 13, 2026

Unity Catalog now generally offers three complementary data-governance capabilities: Attribute-Based Access Control (ABAC) policies for row filtering and column masking, Governed Tags, and agentic Data Classification. They address per-object access rules, coordination gaps, and manual detection by letting governance teams define tag-based policies once, automatically classify sensitive data, and protect matching objects across catalogs and schemas. Governed tags provide an account-level vocabulary inherited across catalogs, schemas, tables, and columns, while ABAC applies row filters and column masks using tag-based conditions. Classification uses built-in compliance classifiers, custom classifiers, metadata, pattern recognition, and large language models, with human-in-the-loop validation and false-positive exclusions. General availability adds 10x larger policy limits, support for 10,000+ policies per metastore, lifecycle management through SQL, APIs, UI, and Terraform, expanded compliance coverage, and custom classifiers in beta.


### [The Convergence of Open Table Formats and Open Catalogs: Catalog Commits is Generally Available](https://yomu.fyi/post/the-convergence-of-open-table-formats-and-open-catalogs-catalog-commit.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Benjamin Mathew, Michelle Leon, Lukas Rupprecht, Ryan Johnson
- Published: May 12, 2026

Catalog Commits is generally available for Unity Catalog managed Delta tables, aligning Delta with Iceberg’s catalog-oriented model and making the catalog responsible for table discovery, access, and latest table state. The change addresses three coordination problems: metadata “split brain” when engines write directly to storage, fragmented multi-engine governance, and the historical inability to coordinate atomic writes across multiple tables. With Catalog Commits enabled, Unity Catalog brokers table access through standardized APIs, keeping catalog and table state synchronized and enabling consistent authorization, holistic auditability, automated optimizations, and multi-statement, multi-table ACID transactions on Databricks. The release supports Databricks products and engines including Delta Spark, Delta Flink, Starburst Trino, DuckDB, and StreamNative, while Delta Kernel provides a shared path for connector support.


### [Addressing HR's widening capacity gap with AI](https://yomu.fyi/post/addressing-hr-s-widening-capacity-gap-with-ai.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Paurul Yadav, Soumya Dash, Bryan Smith
- Published: May 8, 2026

HR teams face a widening capacity gap as strategic expectations, complex employee issues, workforce volatility, skills shortages, and demands for personalized support collide with largely unchanged headcount and tools. The article presents AI transformation as an incremental journey: first establish a secure Employee 360 from structured and unstructured enterprise data, then build reusable workforce insights, augment workflows with human oversight, and progress toward broader transformation. It emphasizes data governance, including access controls, auditing, quality, standardization, and reliable interpretations, while noting that trust has limited AI’s business impact so far. MathCo and Databricks support this roadmap through NucliOS, whose Data Studio, AI Studio, and Decision Studio environments connect governed data, explainable models, feedback loops, and decision applications; Databricks supplies the lakehouse foundation, lineage, quality checks, and privacy-compliant access.


### [MCP Marketplace brings real-time intelligence to agentic applications](https://yomu.fyi/post/mcp-marketplace-brings-real-time-intelligence-to-agentic-applications.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Roman Ostrovski, Harish Gaur, Antoine Amend
- Published: May 8, 2026

The MCP Marketplace connects agentic applications with real-time external intelligence alongside enterprise data. The problem appears in use cases such as loan approval, where historical internal records omit market conditions, updated credit signals, property changes, and competitor activity, making manual research a bottleneck. Databricks Marketplace provides governed access to MCP servers from You.com, Moody’s, and Cotality, while Unity Catalog authenticates connections and tracks access and lineage; Lakebase stores state, decisions, and audit trails across multi-step workflows. Examples show agents combining internal data with web research, credit ratings and sector outlooks, or property-resolution and mortgage signals before surfacing decisions for human review, including a commercial-loan flow with recorded sources, timestamps, and approver.


### [Real-world evidence for medical affairs: who can actually use it?](https://yomu.fyi/post/real-world-evidence-for-medical-affairs-who-can-actually-use-it.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Adam Crown
- Published: May 8, 2026

Real-World Evidence (RWE) is increasingly sought by payers, providers, and regulators, but the post distinguishes it from abundant Real-World Data (RWD), which must undergo rigorous study design, analysis, and interpretation to become credible evidence. It identifies four Medical Affairs use cases: regulatory submissions and commitments, payer formulary discussions, rapid HCP scientific exchange, and internal pipeline or portfolio decisions. The stated operational problem is that teams with claims, EHR, registry, and other assets often lack the fluency or capacity to answer complex questions within competitive and regulatory timelines. Databricks Genie is presented as a natural-language interface that can query unified RWE assets, including a treatment-initiation and 12-month-persistence example that reportedly surfaces in seconds rather than requiring several days of data-science work. The post also describes governance, logged and attributable requests, treatment-pathway awareness, and access controls for MSL support.


### [Public health intelligence shouldn't require a data scientist](https://yomu.fyi/post/public-health-intelligence-shouldn-t-require-a-data-scientist.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Kacey Hertan
- Published: May 7, 2026

State, tribal, local, and territorial (STLT) health agencies manage data across surveillance, vital records, Medicaid, WIC, and emergency preparedness systems, but those systems are fragmented and difficult to query together. That limits rapid public health intelligence: questions linking emergency-department visits with pharmacy dispensing, school absenteeism, vaccination, demographic, or geographic data can require epidemiologists to assemble manual queries over weeks, even when decisions require answers within hours. The post presents Databricks Genie as a natural-language interface for querying this environment, backed by a Databricks engine that handles petabyte-scale datasets across real-time streams and historical records. It describes cross-program synthesis, Unity Catalog row- and column-level access controls, HIPAA-compliant governance, traceability to the underlying query, and validation controlled by health experts. Examples include county-level influenza-like illness trends overlaid with vaccination coverage and identifying counties with high opioid overdose rates and low treatment utilization; Genie is described as available today.


### [Mean time to detect is a data access problem](https://yomu.fyi/post/mean-time-to-detect-is-a-data-access-problem.md)
- Company: [Databricks](https://yomu.fyi/company/databricks.md)
- Author: Taylor Kain
- Published: May 7, 2026

Security operations centers measure MTTD, MTTR, false-positive rates, and analyst utilization, yet investigations often stall because analysts assemble evidence across fragmented systems. A single alert may require separate queries for logs, identity records, asset information, prior alerts, and cross-source timelines, making the analyst the integration layer and creating an MTTI bottleneck. The post presents Lakewatch with Databricks Genie as an agentic interface powered by Anthropic Claude models: analysts ask natural-language questions while autonomous agents hunt, summarize, correlate, and reconstruct timelines across security, IT, and business data. It argues that this architecture can reduce investigation work from manual, multi-system workflows to answers in seconds, while retaining analyst-level access controls and governed data access as exploit time has shrunk to 1.3 days.


[Older posts](https://yomu.fyi/topic/data-governance/page/2.md)
