# Engineering at scale, with purpose.

[Thumbtack](https://yomu.fyi/company/thumbtack) · Thumbtack People Team · Feb 5, 2026

**Type:** Explainer

## Summary

Senior software engineer Brett Shouse outlines site reliability engineering initiatives and infrastructure modernization efforts underway at Thumbtack. The engineering organization is currently leading an operational project to migrate multiple self-hosted observability services onto a single, unified SaaS platform. This architectural transition consolidates application logs, distributed traces, and system metrics into one accessible view tailored for engineers, customer support staff, and company executives. Moving away from legacy self-hosted monitoring systems reduces systems administration overhead, mitigates alert fatigue, and lowers direct operational infrastructure costs. Concurrently, the reliability team tackles technical debt accumulated from rapid organizational growth by establishing structured incident response processes and automating repetitive operational toil.

## Context

Thumbtack faced technical debt accumulated from rapid scaling, as well as ongoing systems administration overhead, application support burdens, and alert fatigue caused by operating multiple self-hosted observability services.

## Approach / What changed

Site reliability engineering is migrating the self-hosted observability services into a unified SaaS platform, consolidating logs, traces, and metrics into a single view across engineering, customer support, and leadership.

## Takeaways

- Migrating self-hosted observability services to a unified SaaS platform lowers infrastructure costs and reduces ongoing systems administration, application support, and alert fatigue.
- Unifying logs, traces, and metrics into a single platform provides an accessible operational view across multiple roles, including engineers, customer support, and executives.
- Addressing technical debt from rapid company growth involves formalizing incident response practices to enable faster problem resolution and automating operational toil.

**Tags:** [Incident Response](https://yomu.fyi/topic/incident-response), [Migrations](https://yomu.fyi/topic/migration), [Observability](https://yomu.fyi/topic/observability), [Reliability](https://yomu.fyi/topic/reliability)

[Read original post](https://medium.com/thumbtack-engineering/engineering-at-scale-with-purpose-f36aa16db839)
