Loading…
You're Spending Too Much on AI. You're Also Using Too Little.
Anand Kuchibotla, Kedar Thakkar, Rahul Sengottuvelu
- Source
- Ramp
- Published
- Added to Yomu
Summary
The post argues that a large AI bill does not show excessive use: companies can overspend on routine work while using too little AI where advanced models could create value. It proposes measuring work in atomic tasks—such as invoices coded or pull requests reviewed—rather than tokens, with cost defined by tasks attempted and value by successful tasks. The operating model uses defaults that pair routine work with the cheapest model meeting quality benchmarks, medium reasoning, and flexible latency, while escalating ambiguous, high-stakes work to frontier models at higher effort. It also recommends attributing spend by provider, product, team, and workflow, finding concentrated costs, benchmarking repeated tasks, repricing after model releases, and centralizing controls in one gateway. The conclusion is that efficiency should be treated as an engineering achievement, so cheaper routine execution funds ambition on rare tasks where extra intelligence can materially change outcomes.
Context
Companies face two linked problems: they overspend on AI used for routine work while underusing AI for tasks where advanced models could deliver substantial value. Provider defaults, weak spend attribution, and incentives that reward outcomes without considering efficiency make waste difficult to identify and control.
Approach / What changed
Measure AI economics in successful units of work rather than tokens; attribute costs by provider, product, team, and workflow; benchmark repeated tasks; set cheaper defaults for model type, reasoning effort, and latency; deliberately escalate difficult or high-stakes work; reprice workloads after model releases; and centralize controls in one gateway.
Takeaways
- Tokens measure usage but do not show business value; atomic tasks such as invoices coded, tickets resolved, or pull requests reviewed provide more legible units for pricing and ROI analysis.
- Routine work should default to the cheapest model that clears quality benchmarks, medium reasoning, and flexible latency; frontier models should be reserved for novel, ambiguous, or high-stakes problems.
- AI cost management should include spend attribution, workflow-level concentration analysis, benchmarks, post-release repricing, and a culture that recognizes efficiency as an engineering achievement.