Loading…
A Guide to Running an Engineering Program
2023-10-18
- Source
- Shopify
- Published
- Added to Yomu
Summary
Shopify describes a playbook for running large engineering programs as platform complexity grows across products, tooling, and architecture. The framework begins with a Program Plan covering the problem statement, objectives, guiding principles, definition of done, risks, mitigation, staffing, scope, and a path to completion. Execution uses six-week cycles that revisit unfinished goals, regressions, emerging risks, and the planned route to completion; the described program spans 200 people, nine sub-organizations, and 92 projects. Supporting rituals include weekly updates, status checks, risk and escalation triage, cycle reviews, retrospectives, RFCs, and performance testing with Lua scripts orchestrated by Genghis. The post concludes that engineering program management must adapt to each organization and that a combination of scheduled and ad hoc rituals helped Shopify pursue its goals.
Context
Shopify’s rapid development, growing commerce surface area, and evolving architecture have increased platform complexity. The company has also taken on large engineering programs spanning multiple areas, requiring alignment among stakeholders, workstreams, and senior leadership across organizations.
Approach / What changed
Define a Program Plan with aligned scope, duration, staffing, outcomes, guiding principles, a program- and workstream-level definition of done, risks and mitigations, and a path to completion. Execute it through six-week cycles and a combination of weekly, cycle-based, and ad hoc rituals, including updates, triage, retrospectives, RFCs, and performance testing. Performance tests use configured shops, Lua scripts, and the internal Genghis tooling.
Takeaways
- The definition of done includes team contribution checklists, a performance baseline, a resiliency plan and gameday, internal documentation, and external documentation.
- Each six-week cycle incorporates unachieved goals or regressions, newly identified risks, and goals expected from the program’s path to completion.
- Teams validate component performance against target Service Level Indicators, then can fold passing components into end-to-end tests covering happy paths and maximum system complexity.