Loading…
How Shopify Manages Petabyte Scale MySQL Backup and Restore
2023-10-18
- Source
- Shopify
- Published
- Added to Yomu
Summary
Shopify manages a petabyte-scale fleet of MySQL servers across replica-sets, or shards, in three Google Cloud Platform regions, where file-based Xtrabackup archives made shard restores take more than six hours. To reduce that Recovery Time Objective, the team built Kubernetes CronJobs around the Compute Engine Persistent Disk snapshot API, scheduling snapshots every 15 minutes while accounting for regions, zones, instance roles, and MySQL consistency variables. Persistent Disk snapshots take about 20 minutes initially and typically under 10 minutes incrementally; restoring the latest snapshot to a new volume and starting MySQL, including InnoDB instance recovery and replication-lag recovery, typically brings RTO below 30 minutes. Retention tooling keeps the latest two snapshots plus policy-defined dailies and weeklies, while verification checks startup, GTID auto-positioning, and InnoDB page corruption; offsite tooling compresses, encrypts, and transfers exported data to another provider.
Context
Shopify’s petabyte-scale MySQL fleet required a faster backup and restore system. File-based Xtrabackup backups archived in Google Cloud Storage made shard restores take more than six hours, creating a high Recovery Time Objective and slowing replica rebuilding and read-scaling operations.
Approach / What changed
The team built Kubernetes-managed tooling around Google Cloud Compute Engine Persistent Disk snapshots. CronJobs create snapshots every 15 minutes, retention jobs enforce regional policies, restore workflows create disks from recent snapshots, and verification jobs test restored databases. Separate tooling compresses, encrypts, and transfers exported snapshot data to an offsite provider.
Takeaways
- Initial multi-terabyte Persistent Disk snapshots took around 20 minutes, while incremental snapshots typically took less than 10 minutes.
- Retention automation deletes thousands of daily snapshots and keeps the latest two per shard alongside policy-defined dailies, weeklies, and other copies in each region.
- Verification restores retained daily snapshots and tests MySQL startup, GTID auto-positioning for replication, and InnoDB page-level corruption; it covers more than a petabyte each day.