Loading…
Adventures in Garbage Collection: Improving GC Performance in our Massive Monolith
2023-10-18
- Source
- Shopify
- Published
- Added to Yomu
Summary
Shopify investigated high Ruby garbage-collection latency in its monolith, where minor marks took about 3 ms but major marks could exceed four seconds and some requests spent multiple seconds in GC. Logging showed that roughly 75% of major marks were triggered by oldmalloc and nearly 25% by shady, while nofree became dominant after oldmalloc was addressed. The team used production experiments, improved metrics, changed GC environment settings, increased initial heap slots, and added out-of-band collection during request processing. A 128 GB oldmalloc threshold reduced major marking by about 20%; later out-of-band collection cut the P99.99 and P99.9 GC-time tail by almost tenfold and largely eliminated median-request pauses. The resulting configuration, with collection frequency adapting from every 128 to every 512 requests, was deployed across the fleet after about three weeks.
Context
Ruby garbage collection pauses other execution while it runs, increasing request latency in Shopify's monolith. Major marks were especially costly, reaching more than four seconds, and frequent deployments kept processes in a warm-up phase where the nofree heuristic could trigger additional major collections.
Approach / What changed
The team improved GC logging and metrics, identified trigger reasons, tested changes on portions of the production fleet, and evaluated each result. Changes included setting both oldmalloc limit variables to 128 GB, increasing initial heap slots to reduce nofree triggers, and using out-of-band garbage collection at an adaptive frequency from once per 128 requests to once per 512 requests as processes aged.
Takeaways
- About 75% of major marks were initially attributed to oldmalloc and nearly 25% to shady; after oldmalloc was raised to 128 GB, nofree became the dominant observed trigger.
- Out-of-band garbage collection every 100 requests made major marking during request cycles a fringe event and reduced P99.99 and P99.9 GC time by almost 10x.
- The final adaptive frequency reduced collection overhead for older processes without degrading tail latency, while median GC time per request increased to about 2 ms.