# The Hugo evolution: Engineering Grab's unified, one-click data ingestion platform with Apache Flink

[Grab](https://yomu.fyi/company/grab) · Shuguang Xiang · May 22, 2026

## Summary

Grab's self-service data platform, Hugo, faced significant onboarding friction as streaming pipelines expanded across fragmented systems like Kafka Connect, custom Go applications, and Spark. Engineering teams struggled with cross-platform configuration translations and brittle, manual schema mappings that stretched onboarding over several days. To resolve these bottlenecks, Grab modernized the ingestion architecture by introducing a centralized automation layer powered by Apache Flink and Flink CDC. The updated platform dynamically retrieves Protobuf schemas from Confluent Schema Registry and ingests MySQL binlogs directly into queryable Hive tables without intermediate Kafka hops. This shift dropped pipeline onboarding times to roughly six minutes for Kafka and three minutes for MySQL CDC, driving more pipeline adoptions in one year than in the previous five.

## Takeaways

- Adopting Flink CDC eliminated intermediate Kafka hops and manually maintained Go DTOs by streaming MySQL binlogs directly into Hive tables with Spark compaction.
- The modernized Kafka ingestion pipeline dynamically queries Confluent Schema Registry on startup, removing the need to hardcode Protobuf-to-Avro mappings in application source code.
- Integrated onboarding guardrails proactively validate database binlog settings, topic ownership, and destination table naming before provisioning begins, cutting setup times to minutes.

**Tags:** [Data Pipelines](https://yomu.fyi/topic/data-pipelines), [Developer Experience](https://yomu.fyi/topic/developer-experience), [Kafka](https://yomu.fyi/topic/kafka), [MySQL](https://yomu.fyi/topic/mysql), [Streaming](https://yomu.fyi/topic/streaming)

[Read original post](https://engineering.grab.com/one-click-data-ingestion-platform-with-apache-flink)
