# Optimally Scaling Kafka Consumer Applications

[Grab](https://yomu.fyi/company/grab) · Shubham Badkur · Oct 13, 2020

**Type:** Problem & solution

## Summary

Grab's Coban platform runs Golang-based stream processing pipelines on Kubernetes, servicing roughly 400 billion events weekly from Kafka. The initial Horizontal Pod Autoscaler setup caused resource waste and uneven load distribution across Kafka partitions during scale-in and scale-out events. To resolve this, Grab moved to a fixed pod count matching the topic's partition count and adopted Vertical Pod Autoscaling, reducing resource usage versus requests by approximately 45%. The team also introduced Kubernetes priority classes to segment latency-sensitive workloads onto On-Demand nodes and non-critical jobs onto Spot instances. Additionally, overprovisioning via low-priority placeholder pods managed by Cluster Proportional Autoscaler enabled rapid pod rescheduling and reduced deployment delays.

## Context

Horizontal Pod Autoscaling on Grab's stream processing platform led to resource wastage from conservative provisioning for peak traffic and resulted in uneven Kafka partition assignment across scaling pods.

## Approach / What changed

Grab fixed the number of pipeline pods to match Kafka partition counts retrieved at deployment, replaced HPA with Vertical Pod Autoscaling, separated workloads onto Spot and On-Demand nodes using priority classes and node affinity, and added overprovisioned placeholder pods managed by Cluster Proportional Autoscaler.

## Takeaways

- Aligning the number of consumer pods directly with the Kafka topic partition count prevented uneven partition distribution during scaling events.
- Replacing Horizontal Pod Autoscaling with Vertical Pod Autoscaling on fixed-size consumer deployments yielded an approximate 45% reduction in total resource usage relative to requested capacity.
- Using low-priority placeholder pods scaled by Cluster Proportional Autoscaler allowed immediate preemption and rescheduling when Spot instances were terminated.

**Tags:** [AWS](https://yomu.fyi/topic/aws), [Kafka](https://yomu.fyi/topic/kafka), [Kubernetes](https://yomu.fyi/topic/kubernetes), [Performance](https://yomu.fyi/topic/performance), [Streaming](https://yomu.fyi/topic/streaming)

[Read original post](https://engineering.grab.com/optimally-scaling-kafka-consumer-applications)
