Cut AI Agent Costs with TokenOps
Gareth Bland explains TokenOps: a practical way to control AI agent spend by focusing token budget where it creates real value.
Overview
Why AI agent costs grow faster than the work
- As an agent takes more turns, context accumulates, which increases the amount of input that must be sent repeatedly.
- Uncached input can cause costs to rise on a quadratic curve.
- Example given: if an agent takes 4× as many turns, it can cost 16× as much.
Why output tokens can be the expensive part
- Gareth highlights that reasoning can make output tokens the more expensive component compared to input tokens.
Using caching to reduce costs
- Engineering the agent’s context to use cached input can significantly reduce spend.
- The example claim in the description: caching can bring the cost curve down to a tenth of the cost.
Measuring “agentic efficiency”
Gareth proposes measuring agentic efficiency using three ratios:
- Autonomous completion rate
- Token efficiency ratio
- Value generation ratio
Deciding where the next token dollar belongs
- TokenOps is framed as deciding where token spend will “move the needle most.”
- An S-curve model is used to interpret investment levels:
- Underinvesting
- Getting real leverage
- Gold-plating features nobody asked for
Resources
- Unit Economics of Agentic AI: https://aka.ms/tokenomicon-2026
- Maximize ROI from AI: https://azure.microsoft.com/solutions/maximize-roi-from-ai
- More Azure resources: https://azure.com/AzureEssentials
Connect
- Gareth Bland: https://www.linkedin.com/in/gareth-bland/