4X the Turns, 16X the Cost
Gareth Bland explains why AI agent costs can scale much faster than the work being done, and how context caching helps control token spend as agents take more turns.
Overview
- The short highlights a cost-scaling problem in agentic AI: increasing the number of agent “turns” can drive token usage and cost up disproportionately.
- Gareth Bland points to caching (specifically, caching context) as a practical technique to reduce repeated token spend and bring costs back down.
- The clip is taken from a longer discussion on TokenOps: https://youtu.be/B4yovi_znxE