Key Takeaways
- AI agents now consume far more tokens than human users on the OpenRouter platform—about 5× as much, with the gap expected to widen to 10× or higher.
- Over 85 % of agent‑generated tokens come from cached prompts, meaning the same information is repeatedly re‑read rather than newly generated.
- Agent token usage has grown 14× since the February crossover point, while human usage has risen only 2.8×; a mixed “agent‑human” category grew 4.7×.
- The surge in cached token storage is driving demand for high‑bandwidth memory (HBM), with KV caches already beginning to exceed current GPU HBM capacity.
- Industry surveys show AI‑agent adoption accelerating: 40 % of large organizations reported scaling agents in McKinsey’s 2026 State of AI survey, up from 27 % a year earlier.
- Memory makers are prioritizing HBM for AI data centers, and analysts warn RAM and storage shortages could worsen in 2027‑2028 as agent workloads continue to expand.
Overview of Agent vs. Human Token Usage
Futurum Group CEO Daniel Newman highlighted a striking imbalance in AI consumption when he wrote on X, “AI is currently used by AI 5x more than it is used by humans. That number will accelerate to 10x and then higher and higher.” His claim is backed by an Andreessen Horowitz (a16z) chart derived from OpenRouter data, which shows agents consuming 7.3 trillion tokens versus 1.4 trillion tokens generated by human users as of August—roughly six months after agent usage first surpassed human usage. This disparity illustrates that the bulk of AI compute is now being driven by autonomous systems rather than direct human interaction.
Breakdown of Token Sources and Cached Data
OpenRouter classifies each API key into one of three buckets—agentic, mixed, or human—using a “7‑signal weighted composite score” that factors in tool‑call rate, turn count, gap timing, and other metrics. According to the platform’s head of insights, Peter Walker, the seven‑day average token usage split reveals that agents are using 14× more tokens while human usage is up only 2.8× since the February crossover. The mixed category, which likely captures hybrid agent‑human workflows, grew 4.7× over the same period. Crucially, a16z noted that more than 85 % of agent tokens originate from cached prompts, meaning the same context is repeatedly re‑read rather than newly generated. As one journalist put it, “Cached tokens also account for nearly all of the relative growth in token usage… they cost far less than processing a prompt from scratch, but they still have to be held in memory.”
Implications for Memory Hardware
Because cached prompts must reside in memory, the explosion of agent token reuse is translating into heightened demand for high‑bandwidth memory (HBM). Models store this reusable context in the key‑value (KV) cache, and a16z warned that “the KV cache is outgrowing GPU HBM capacity.” This observation was echoed in internal logs from a call‑center consultancy that tested DeepSeek on rented Nvidia H200 GPUs; the consultancy reported that, in its own agents’ September usage on Claude Code, “96 % of all input was re‑reading old conversation.” Such statistics suggest that while raw token counts may overstate the direct cost of new computation, the underlying hardware pressure—particularly on memory bandwidth and capacity—remains very real and escalating.
Industry Adoption Trends
Beyond OpenRouter, broader market data corroborates the rise of AI agents. McKinsey’s 2026 State of AI survey found that 40 % of respondents from large organizations reported scaling AI agents, up from 27 % the previous year. This upward trajectory indicates that enterprises are moving beyond pilot projects to production‑scale deployments of agentic workflows, further amplifying the token and memory demands highlighted by the OpenRouter metrics. The trend is not isolated to a single platform; similar patterns have been observed across multiple cloud providers and AI‑focused startups, suggesting a systemic shift toward agent‑centric AI architectures.
Future Outlook and Potential Bottlenecks
If Newman’s projection holds—that the 5× agent‑to‑human token ratio will climb to 10×, 20×, 30× and beyond—the strain on memory subsystem resources will intensify. Micron has already forecast that RAM and storage shortages will worsen in 2027 and 2028, with customers likely to pay premiums as memory makers prioritize HBM for AI data centers. Consequently, PC buyers and other end‑users could find themselves competing directly with massive agent workloads for limited memory supplies, potentially driving up prices and constraining availability for traditional computing tasks. The scenario underscores a critical inflection point: as AI agents become the dominant consumers of compute, the industry must innovate not only in algorithmic efficiency but also in memory technology and system architecture to keep pace with exponential growth.
https://www.tomshardware.com/tech-industry/artificial-intelligence/futurum-ceo-says-agents-use-ai-5x-more-than-humans-number-will-eventually-hit-10x-but-agents-are-mostly-rereading-what-theyve-already-seen

