Loading…
Announcing the Databricks storage ecosystem: Governing the enterprise data estate, wherever it lives
Rupal Jain, Denis Dubeau
- Source
- Databricks
- Published
- Added to Yomu
Summary
Databricks announces a Software-Defined Storage (SDS) Ecosystem for governing enterprise data across on-premises, private-cloud, and edge environments without requiring migration. It responds to sovereignty and regulatory constraints, data-gravity economics, latency requirements, and the need to unlock AI value in backup, archive, and other previously inaccessible data. Storage partners implement the open-source OpenSharing protocol, expose an endpoint, and connect it to Unity Catalog so Databricks Serverless Compute can access governed data where it resides. The integrations provide a unified catalog and support Serverless Compute, Genie, AgentBricks, and model training with zero data movement or duplication; MinIO is generally available, while Everpure, Qumulo, and VAST Data are in preview stages. Databricks also says commitments are in place from six additional providers and that Volumes APIs are being developed to extend OpenSharing to unstructured files for generative-AI workloads.
Context
Enterprises with sovereignty requirements, extreme data volumes, high egress costs, latency-sensitive workloads, or large on-premises data estates cannot move all of their data to the cloud. Databricks describes a shift from “Migrate Everything” to “Govern Everything,” driven partly by demand from hundreds of customers for on-premises and hybrid storage connectivity to Unity Catalog.
Approach / What changed
The SDS Ecosystem uses storage-provider implementations of the open-source OpenSharing protocol. A provider exposes an OpenSharing endpoint, which connects to Unity Catalog and enables Databricks Serverless Compute to access governed on-premises data without migration, replication, or duplication. Databricks is also developing Volumes APIs for direct access to unstructured files.
Takeaways
- MinIO AIStor is generally available and supports governed access to live on-premises Apache Iceberg and Delta tables through OpenSharing.
- Everpure, Qumulo, and VAST Data have announced OpenSharing integrations at private-preview stages, with Qumulo scheduled for July 2026 and VAST Data for August 2026.
- Databricks has commitments from Cohesity, Commvault, HPE, NetApp, Nutanix, and Rubrik, while extending OpenSharing toward unstructured data for generative-AI workloads.