Loading…
Thanos
1 posts about Thanos. Every summary links to the original.
10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks
Databricks’ monitoring infrastructure now tracks 5 billion active timeseries in real time and ingests more than 10 trillion samples daily, exposing scalability, reliability, cost, and operability limits in its older stack. The replacement, Pantheon, is a fork of CNCF Thanos deployed across more than 160 instances in about 70 regions and three cloud providers; tiered storage, differentiated memory retention, isolated replicated Receive groups, multitenancy, and a custom control plane support automated scaling and recovery. Pantheon’s largest instance holds about 300 million in-memory timeseries and handles nearly 1,000 PromQL queries per second, while migration reduced annual cloud costs by millions and monitoring downtime by roughly five times. For high-cardinality troubleshooting, Hydra preserves raw metrics in Delta tables, exposes them through Grafana and SQL, and unifies metric semantics across aggregated and raw paths, with freshness improvements planned.
David Yuan, Yi Jin, Karan Bavishi, HC Zhu, Joey Beyda