Loading…
This Rocket Ain't Stopping - Achieving Zero Downtime for Rails to Golang API Migration
GrabLian Yuanlin
Summary
Grab transitioned its public passenger app APIs from a legacy Rails application to a Golang service-oriented architecture to consolidate its codebase and engineering teams. Initial attempts to proxy traffic through a cloned Rails server via gRPC were abandoned after encountering TCP load imbalances during autoscaling events and memory leaks in the gRPC Ruby gem. The team pivoted to direct logic migration, porting Ruby logic directly into Go while decomposing modules into standalone services. Verification relied on log-based load testing and live shadow testing, where write operations were safely validated using mock data access layers that evaluated expected database outcomes. Production rollout progressed endpoint-by-endpoint using requests-per-second traffic throttling and prewarmed AWS Elastic Load Balancers before executing the final DNS switch.
Context
Grab was transitioning from a Rails and NodeJS stack to a full Golang Service Oriented Architecture and needed to migrate live public passenger app APIs from an existing Rails app to a new Go server cluster without downtime.
Approach / What changed
The team migrated endpoints individually across four phases—logic migration, log-based load testing, shadow testing with mock data access layers for write requests, and progressive rollout using requests-per-second throttling and prewarmed ELBs before cutting over DNS.
Takeaways
- Initial proxying between Go and a Rails clone via gRPC was abandoned due to TCP persistent connection load imbalances during autoscaling events and a memory leak in the gRPC RubyGem.
- Shadow testing non-idempotent write requests (PUT/POST/DELETE) was achieved without duplicate writes by wrapping data access objects in mock code that generated and verified expected database rows.
- Throttling shadow and rollout traffic by percentage caused ELB request drops on high-traffic endpoints, prompting a switch to explicit requests-per-second (RPS) rate limits and ELB prewarming.
Related reading
Grab ·
Go module proxy at Grab
Grab's 69.3 GiB multi-module Go monorepo caused commands like go get to take over 18 minutes as Git repeatedly traversed commit history, downloaded large worktrees, and overloaded their GitLab VCS infrastructure. To bypass direct VCS queries without losing automatic updates for external repositories, the team deployed the Athens Go module proxy configured in fallback network mode. They used the GOVCS environment variable to disable Git access specifically for the monorepo path, forcing Athens to fall back to its internal object storage when resolving monorepo modules. A dedicated CI pipeline pre-populates and refreshes the Athens cache whenever new monorepo modules are released. This setup reduced monorepo go get execution times to approximately 12 seconds and allowed a 70% scale-down of the Athens proxy cluster.
Jerry NgGrab ·
Taming the monorepo beast: Our journey to a leaner, faster GitLab repo
Grab's decade-old Go monorepo grew to 12.7 million commits and 250GB of Git data, causing Gitaly replication delays of up to four minutes that routed all read traffic exclusively to the primary node and slowed developer operations. After staging tests proved that shallow history reduced replication lag from hundreds of seconds to under three seconds, standard rewriting tools like git filter-repo and git rebase failed due to complex merge histories and repository scale. To overcome runner memory limits and lengthy git garbage collection cycles, the engineering team implemented a custom two-phase migration script. The script selectively migrated 2,000+ critical dependency tags and one month of recent history, flattening merge commits, embedding legacy hashes for traceability, and reducing total commit volume by 99.9%.
Nagendra GangwarGrab ·
Turbocharging GrabUnlimited with Temporal
GrabUnlimited experienced scaling bottlenecks, corrupted membership states, and elevated production incidents after its subscriber base grew by over 1000%. The original architecture relied on Amazon SQS state machines, 5-minute Redis locks, and daily batch cron jobs that overwhelmed the database and lacked granular idempotency during upstream retries. To eliminate these failure modes, the engineering team migrated the core membership lifecycle to Temporal's workflow orchestration engine. Replacing batch cron jobs with Temporal Timers distributed renewal operations throughout the day, while matching workflow IDs prevented race conditions between renewals and cancellations. This architectural transition resolved database bottlenecks and yielded an 80% reduction in open production incidents.
Michel ParrenoNetflix ·
Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned
Engineers at Netflix required a unified, real-time view of service dependencies to navigate distributed architecture and improve incident troubleshooting. Traditional batch systems introduced stale data, so the team created a streaming-first platform backed by reactive streams and backpressure handling to ingest flow records from multi-region Kafka streams and Server-Sent Events without data loss. The architecture partitions data into physically separate graph and columnar storage layers covering eBPF network flows, IPC metrics, and distributed traces. Network flow ingestion relies on a three-stage distributed aggregation pipeline using consistent hashing to resolve network intermediaries into logical application connections. The resulting production system serves time-travel and topology queries with sub-second latency while continuously updating dependency views.
Netflix Technology Blog