Loading…
Latest reads
The engineering internet, summarised so you can actually read it.
Grab ·
Being a Principal Engineer at Grab
Grab's rapid growth resulted in roughly 350 microservices powering its superapp, creating the need for defined individual contributor career milestones. At Grab, a principal engineer oversees the architecture of an entire Tech Family, a sub-organisation containing over 50 engineers and 20 or more microservices. Responsibilities include translating broad, ambiguous problems into concrete projects, managing technical debt, and aligning multiple engineering teams across global R&D centres. The role demands continuous technical leadership through RFC design reviews, cross-functional communication, mentorship, and self-directed prioritization alongside engineering leadership. Ultimately, principal engineers amplify engineering quality and operational stability without directly managing people.
Roman AtachiantsGrab ·
Data First, SLA Always
Grab's Data Engineering team transitioned from periodic batch ETL ingestion to a real-time change data capture architecture called Trailblazer after dataset sizes exceeded the petabyte mark. The previous system caused severe JDBC timeouts and heavy CPU loads when executing chunked or full-scan queries on unindexed upstream MySQL tables. To solve this, MySQL binary logs are captured via Debezium on Kafka Connect, buffered in Kafka, and ingested into a data lake using Spark Structured Streaming. Checkpoints are decoupled from local storage and persisted in a Redis cluster to simplify ingestion offset overrides and handle ephemeral compute clusters. The system incorporates extensive health monitoring across Airflow, Datadog, and custom services to maintain stream liveliness and avoid Kafka retention breaches.
Johan KokGrab ·
Save Your Place with Grab!
Grab introduced Saved Places across Southeast Asia to eliminate the friction of repeatedly typing addresses and prevent selection errors between similarly named locations. Data analysis across transport and food orders revealed that consumers consistently visit only five to seven unique locations and order food to one or two addresses. To streamline repeat bookings, cross-functional teams designed a feature allowing users to bookmark locations under custom labels such as Home and Work. Following usability testing and release, more than 14 million users stored nearly 45 million addresses across the platform. Platform metrics also showed that while office destinations clustered in central districts across major cities, residential distributions varied significantly between markets like Singapore and Jakarta.
Summit SauravGrab ·
No More Forgetting to Input ERP Charges - Hello Automated ERP!
Grab launched an automated Electronic Road Pricing (ERP) fare calculation feature in Singapore to eliminate the need for driver-partners to manually track gantries and enter toll charges. Because Singapore gantries frequently adjust fares based on time and road conditions, manual entry often caused driver errors and revenue loss. Grab solved this by mapping precise geographical coordinates for every toll gate using satellite imagery and open data, matching frequent driver GPS pings against road layers and gantry locations. The engineering and operations teams also built an internal ERP Workflow tool to map ride trajectories and resolve driver dispute feedback within an average of one day. Following its rollout in Singapore, Grab began testing and planning regional expansion to Indonesia, Thailand, Malaysia, and the Philippines.
Garvee GargGrab ·
How We Built a Logging Stack at Grab
Grab needed a scalable logging platform to replace slow, fragmented systems that hindered debugging across their growing service fleet. Generating 25TB of daily logs, the team built a horizontally scalable Elasticsearch cluster configured via Ansible and monitored with Datadog. Although the initial proof of concept assigned all node roles (ingest, coordinator, master, and data) to every machine, operating at scale introduced major challenges with JVM heap exhaustion and cluster stability. The team resolved memory pressure and performance bottlenecks by tuning circuit breakers, lowering field data cache limits, adjusting shard allocations based on segment memory, and disabling translog compression during shard transfers.
Daniel KasenGrab ·
Making Grab’s Everyday App Super
Grab manages an expanding superapp ecosystem comprising ride-hailing, food delivery, payments, and partner content surfaced through the Grab Feed. As content volume grows, the platform risks overwhelming users with irrelevant information. To address this, Grab built a recommendation engine that ranks cards using signals across user profiles, content metadata, and contextual factors such as time and location. The system employs multiple recommendation strategies—including popularity metrics, user favorites, collaborative filtering, habitual patterns, and cross-platform deep embeddings—which are selected or aggregated. Recommendation quality is evaluated via offline metrics like Recall@K and NDCG alongside online engagement experiments.
Justin BoliliaGrab ·
Catwalk: Serving Machine Learning Models at Scale
As machine learning adoption expanded at Grab, individual teams created fragmented model serving solutions that duplicated engineering effort and required data scientists to handle underlying infrastructure. To resolve these inefficiencies, Grab developed Catwalk, a self-service machine learning model serving platform. The system runs TensorFlow Serving containers across a managed Kubernetes cluster integrated with Grab's observability stack. Data scientists deploy or update models simply by saving files using the tf.saved_model API to dedicated Amazon S3 buckets, while Kubernetes automates orchestration, ingress routing, and pod autoscaling. Catwalk abstracts server management away from data scientists, shortens deployment timelines, and provides high availability during model version rollouts.
Nutdanai PhansooksaiGrab ·
React Native in GrabPay
Following the release of the GrabPay Merchant App, Grab adopted React Native inside the Grab Passenger app to maintain a single cross-platform codebase across iOS and Android. Integrating the framework into native applications required establishing bridge communication guidelines and incorporating React Native modules into the Grablet architecture. To streamline network communication, API calls were migrated from axios to native bridges returning promises, which eliminated the need to pass access tokens into JavaScript. The team also established a shared internal library of approximately 20 UI components alongside Redux for state management and react-navigation for routing. Modules like BillPay and Transaction History successfully launched across Southeast Asia while maintaining the performance and feel of native software.
Sushant TiwariGrab ·
Connecting the Invisibles to Design Seamless Experiences
Service design operates as connective tissue within complex product ecosystems by bridging digital touchpoints, physical operations, and backstage technical workflows. At Grab, placing an order on GrabFood requires coordinating driver allocation, customer support paths, and long-term data storage rather than simply transmitting information to merchants. Focusing exclusively on singular features risks breaking broader network dependencies when modifications ripple into other operational systems. Grab addresses these interdependencies through participatory design processes and visual mapping across cross-functional teams. This holistic framework evaluates whether systemic issues, such as inaccurate restaurant operating hours, are best resolved through in-app feature changes or operational adjustments.
Stephanie LukitoGrab ·
Tourists on GrabChat!
Grab examined more than 3.7 million tourist messages across Singapore, Malaysia, and Indonesia sent between December 2018 and March 2019 to evaluate passenger communication patterns. The platform deployed in-house translation and prewritten, auto-translated chat templates to bridge language barriers between international riders and local drivers. Analysis showed that bookings utilizing chat templates experienced a 10% higher ride completion rate than those without. Image-sharing features were most heavily used in high-traffic hubs such as airports, shopping malls, and major tourist centers to aid driver location. Passengers also consistently used messaging to clarify luggage capacity, provide identifiable passenger descriptions, and check pet policies.
Lara PuReum YimGrab ·
Bubble Tea Craze on GrabFood!
GrabFood recorded a regional average order growth rate of 3,000% for bubble tea across Southeast Asia in 2018. Individual country growth rates ranged from over 250% in Malaysia to more than 8,500% in Indonesia over their respective tracking periods. The customer base for bubble tea expanded by over 12,000%, supported by a 200% increase in merchant outlets to nearly 4,000 locations representing over 1,500 brands. Southeast Asian consumers ordered an average of four cups per person per month, led by Thailand at six cups and the Philippines at five cups. Order timing concentrated primarily around lunchtime meals and midday afternoon breaks.
Lara PuReum YimGrab ·
Why You Should Organise an Immersion Trip for Your Next Project
Grab Ventures relies on on-the-ground immersion trips to uncover real consumer behaviors and validate hypotheses that desktop research cannot address. Placing cross-functional teams directly into target markets reveals nuanced consumer motivations, such as Indonesian grocery shoppers prioritizing physical product freshness and price sensitivity over mere convenience. Effective immersion trips require pre-fieldwork reconnaissance by local residents, collaborative hypothesis generation across business and tech disciplines, and non-leading question design. Research teams should operate in small groups of two or three alongside experienced local translators while holding structured end-of-day debriefs to synthesize contextual observations. Within the Double Diamond framework, these field insights support the Discover phase and directly feed into design sprint workshops for problem framing.
Sherizan SheikhGrab ·
Preventing Pipeline Calls from Crashing Redis Clusters
A single Redis slave node failure caused Grab's Apollo booking service to suffer an over 95 percent call failure rate for one minute despite running a three-shard cluster with two replicas per partition. Investigation revealed that the service configured Go-Redis to route all read queries exclusively to slave nodes to offload master CPU usage. When the slave node dropped offline, batched HMGET pipeline calls failed completely because the client wrapper treated a single command failure as a failure of the entire pipeline. Furthermore, the Go-Redis client cached cluster topology and only lazily refreshed state every sixty seconds, continuing to direct traffic to the dead replica until the timer expired. Grab addressed this risk by recommending dedicated pipeline clients configured with latency-based routing to allow reads to fall back to responsive master nodes.
Michael CartmellGrab ·
Guiding You Door-to-Door via Our Superapp!
Grab addressed passenger navigation challenges at large Southeast Asian venues such as airports and shopping centers. Satellite signals weaken through concrete and steel, creating GPS inaccuracies that caused the rendezvous distance between passengers and drivers at large venues to exceed twice the average. While introducing Entrances previously mapped over 120,000 green dots to lower rendezvous distances, passengers still struggled to locate specific pickup spots indoors. The company launched Venues, an in-app feature delivering turn-by-turn text and photo directions to designated pickup points. To support this system, operations teams surveyed sites with cameras and scanners to capture landmarks, after which in-house teams masked faces and vehicle license plates.
Neeraj MishraGrab ·
Loki, a Dynamic Mock Server for HTTP/TCP Testing
Grab built Loki, a dynamic mock server written in Golang that simulates backend services on local developer machines and CI pipelines. Mobile app testing previously suffered from heavy dependencies on complex, brittle staging environments and interconnected services communicating over HTTP, HTTPS, and TCP. Loki handles both HTTP and TCP traffic on distinct ports while exposing a unified RESTful API to manage test expectations. It provides runtime flexibility through sandboxed JavaScript execution, configurable request sequence ordering, and an in-memory cron scheduler for TCP push messages. Adopting Loki decoupled mobile releases from staging stability, improving delivery cycles and enabling automated UI testing with Espresso and XCUITest.
Thuy NguyenGrab ·
How We Harnessed the Wisdom of Crowds to Improve Restaurant Location Accuracy
Grab discovered that abnormally short driver wait times often indicated restaurants registered at incorrect coordinates due to moves or onboarding errors. To fix this, Grab used driver-partner GPS pings, timestamps, and order status updates to infer true food collection locations. The system cleans the data by filtering low-quality GPS pings and isolating the longest temporal streak a driver spends within a predefined radius of the venue. Clusters of inferred pick-up points are then ranked by order volume, the proportion of off-target pick-ups, and median distance errors before routing to mapping operations for verification. This periodic correction workflow achieved a fivefold reduction in order cancellations caused by unfound merchant locations.
Pravin KakarGrab ·
Designing Resilient Systems Beyond Retries (Part 3): Architecture Patterns and Chaos Engineering
Building resilient systems requires architectural safeguards and proactive testing beyond basic retries and circuit breakers. Architectural patterns such as idempotency keys enable safe retries without creating inconsistent state during failures. Asynchronous responses and deferrable work isolate services from downstream dependency latency and errors, though they can conflict with the fail-fast principle. To validate system behavior under stress, chaos engineering introduces intentional failures in production to test hypotheses against a defined steady state. Selectively adopting complementary patterns reduces failure points while avoiding unnecessary architectural complexity.
Michael CartmellGrab ·
Designing Resilient Systems Beyond Retries (Part 2): Bulkheading, Load Balancing, and Fallbacks
Software systems require mechanisms beyond retries to maintain resilience during downstream outages and high traffic. Bulkheading isolates failures across infrastructure, processes, thread pools, and connection limits, preventing a single failing component from degrading an entire system. Load balancing distributes traffic across backend pools via proxies, client-side libraries, lookaside services, or sidecars, often pairing with health checks to eliminate single points of failure. When operations fail unrecoverably, fallback strategies like silent failures, local defaults, stale cache reads, and dedicated backup services enable graceful degradation. Organizations like Grab implement these approaches using internal client-side load balancers backed by etcd, cache fallbacks in microservice frameworks, and redundant core backup services.
Michael CartmellGrab ·
Designing Resilient Systems Beyond Retries (Part 1): Rate-Limiting
Distributed systems that rely exclusively on retries and circuit breakers face severe failure risks, including retry storms and reliance on client-side configuration accuracy. Implementing server-side rate limiting serves as a critical defensive layer to safeguard services across evolving architectures. Throttling thresholds can be layered across per-client, per-endpoint, and server-wide granularities using algorithms such as leaky bucket or sliding windows. While local instance-level limits fail when downstream bottlenecks like databases saturate under horizontal scaling, global rate limiting coordinates traffic enforcement across entire service pools. Centralized rate limiters require asynchronous communication and fallback mechanisms to avoid becoming single points of failure or adding request path latency.
Michael CartmellGrab ·
Context Deadlines and How to Set Them
Microservice architectures operating under heavy network traffic require robust timeout handling to prevent slow or failing dependencies from causing cascading failures across services. Naive, static timeout configurations across a call chain often cause downstream components to waste compute effort on requests that upstream callers have already abandoned. To establish predictable timeout thresholds, engineers can align limits with service-level latency percentiles, pairing P99 limits with median latency retry allowances. Go's context package improves upon static network timeouts by propagating request-scoped deadlines and cancellation signals across service boundaries. This distributed context ensures downstream servers recognize remaining time budgets and terminate unneeded processing immediately when parent deadlines expire or callers manually cancel requests.
Michael Cartmell