Loading…
Ingest semi-structured data faster and more efficiently with Variant - Now Generally Available
Jonathan Brito, Gene Pang, Harsh Motwani
- Source
- Databricks
- Published
- Added to Yomu
Summary
Databricks announces general availability of the Variant data type for ingesting semi-structured JSON, XML, and CSV while avoiding the usual choice between flexible storage and fast, schematized queries. Variant Shredding, also generally available, stores common fields as columns in Parquet files, while Predictive Optimization uses workload and query patterns to identify important fields, collect statistics, and improve file skipping. More than 5,000 teams write Variant, with users executing over 500 million Variant queries monthly across more than 160 TB of data; shredding delivers nearly four-times-faster reads than unshredded Variant and 30-times-faster reads than JSON strings. Auto Loader and Zerobus can ingest Variant into Delta or Iceberg, and future plans include Liquid Clustering by Variant fields, expanded SQL functions, and additional integrations.
Context
Ingesting semi-structured data required choosing between schematizing data for faster queries and storing it as strings for flexibility, while schema changes from upstream applications could force downstream pipeline updates and backfills.
Approach / What changed
Variant enables flexible ingestion into tables, with Variant Shredding storing common fields as Parquet columns. Predictive Optimization uses workload and query patterns to select fields for shredding and statistics collection, improving file skipping. Auto Loader and Zerobus can write Variant data into Delta or Iceberg.
Takeaways
- Variant supports flexible ingestion of events, API JSON payloads, and schemaless database data, including sources such as Kinesis, Event Hub, PostgreSQL, and MongoDB.
- Variant Shredding delivers nearly four-times-faster reads than unshredded Variant and 30-times-faster reads than storing JSON as a string.
- Auto Loader and Zerobus ingest Variant into Delta or Iceberg, allowing clients across the lakehouse to interoperate with the data.