Loading…
The Journey of Deploying Apache Airflow at Grab
GrabChandulal Kavar
Summary
Engineering and data teams across Grab originally operated independent Apache Airflow instances, causing duplicate maintenance overhead and frequent job failures around scaling, logging, and dependency management. To resolve this, a dedicated team developed a centralized orchestration platform that runs isolated, containerized Airflow instances per team on Amazon EKS. The platform categorizes deployments into three size tiers and provisions dedicated Redis brokers, RDS metadata stores, and Vault secret sidecars using Terraform and custom Helm charts. Teams customize container images using shared GitLab CI/CD templates, while worker scaling is handled via Kubernetes Horizontal Pod Autoscalers. Today, the platform runs roughly 20 Airflow instances executing between 1,000 and 60,000 daily jobs per instance.
Context
Multiple teams ran their own Airflow setups without dedicated expertise, leading to resource inefficiency, frequent job failures, and duplicated efforts around scaling, upgrading, and logging.
Approach / What changed
A centralized team built a multi-instance orchestration platform on Amazon EKS using Terraform and Helm, isolating each team to a dedicated Kubernetes namespace with its own RDS, Redis, and customized Docker images built from a shared base image.
Takeaways
- Airflow workers scale dynamically using Kubernetes Horizontal Pod Autoscalers driven by CPU and memory metrics.
- Instances are categorized into small, medium, and large sizing tiers with tailored RDS, Redis, and component resource allocations.
- Airflow 1.10.x lacked native support for Datadog metric tags, prompting internal patching and upstream contributions ahead of Airflow 2.0.0.
Related reading
Grab ·
How Grab Leveraged Performance Marketing Automation to Improve Conversion Rates by 30%
Grab faced operational bottlenecks managing direct-response Google Ads campaigns across thousands of ad groups due to its hyperlocal marketing across Southeast Asian markets. To eliminate the manual burden of tracking and updating ad creatives, the team built CARA, an in-house automation tool deployed on AWS serverless compute. CARA utilizes standardized file naming conventions to map assets to specific campaigns and connects with Google Ads and YouTube APIs to detect and replace low-performing assets. During an experimental rollout across more than 8,000 active ad groups, CARA replaced nearly 2,000 underperforming creatives. The automated asset replacement workflow produced an 18% to 30% increase in clickthrough and conversion rates.
Sc NgGrab ·
7 Fun Facts about Grab’s Driver-Partners in Singapore
Grab analyzed ride-hailing metrics from driver-partners operating in Singapore to identify platform usage trends and driving patterns. Findings indicate that drivers have a 1 in 400 chance of encountering a repeat passenger among the 5.4 million population, with Tampines recording the most pickups and Orchard and Marina Bay serving as top destinations in 2018. Driver behavior data shows that partners with over two years of platform experience routinely start shifts an hour earlier and leverage auto-accept features to minimize idle waiting time. Furthermore, drivers are twice as likely to receive back-to-back ride allocations during evening peak hours, resulting in roughly 50% higher hourly earnings. The dataset also highlights customer satisfaction metrics, showing that shared GrabShare rides achieved an average rating of 4.8 stars.
Lara PuReum YimGrab ·
Customer Support Workforce Routing
Grab replaced its third-party customer support routing software with an in-house workforce routing system for Livechat to gain better priority controls, bespoke configurations, and deeper analytics. The platform separates requests into distinct priority and business queues, using parallel workers that spend varied time slices dequeuing higher-priority issues like safety concerns. To prevent request starvation, workers operate out of sync across queue priority levels while dynamic queue limits cap incoming volume based on agent availability and performance. The system routes requests through an intermediate Agent Group layer, calculating eligibility scores from proficiency and concurrency metrics while managing per-agent locks to prevent over-allocation.
Suman AnandGrab ·
Building a Hyper Self-Service, Distributed Tracing and Feedback System for Rule & Machine Learning (ML) Predictions
Grab's Trust, Identity, Safety, and Security team processes billions of daily rule and machine learning decisions for fraud detection, safety, and identity checks. Earlier logging approaches using plain text Kibana logs and the ActionTrace library lacked structured formats, dynamic entity customization, and fine-grained access controls. To resolve these limitations, the team built Archivist, a centralized tracing, statistics, and feedback system. Archivist ingests events through an SDK into Kafka streams, buffers and routes data into Elasticsearch indices and Amazon S3, and provides a role-based user portal. The platform handles 80 million daily logs across roughly 50 business scenarios, reducing scenario onboarding times from days to minutes.
Warren Zhou