US Army Burned Through a Year of AI Tokens in 38 Days: The Unlimited AI Token Myth
The US Army burned through a year of AI tokens in 38 days. Meta and Uber face the same reality: unlimited AI tokens are never unlimited. What heavy users must know.
The Day the Tokens Ran Out

In May 2026, the US Army’s Chief Information Officer announced that all 3.5 million Army employees would get unlimited AI tokens through the Ask Sage platform. By mid-June, those tokens were gone. A year’s worth of enterprise AI capacity had been consumed in roughly one month.
The email that went out to DEVCOM staff did not mince words: “Although the Army CIO announced in May 2026 that they were offering unlimited tokens, by mid-June the Army CIO pool was exhausted of tokens and had to re-establish limits.”
This is not just a military procurement story. It is the clearest possible illustration of a pattern that is hitting every organization that adopts AI at scale. The phrase “unlimited tokens” is marketing. The reality is finite capacity, shared pools, and hard caps that appear exactly when you least expect them.
What Actually Happened
The Army contracted with Ask Sage, a platform the DOD uses to run multiple LLMs including Google Gemini, Meta Llama, and OpenAI ChatGPT. The enterprise subscription provided access to 100 million tokens per year. For context, a single token in Ask Sage equals roughly 3.7 characters of output.
During Operation Epic Fury, a 38-day military operation in Iran, the Defense Department burned through millions of tokens per day. Combined with rapid adoption across the broader Army workforce, the entire annual token pool was depleted by mid-June.
The Army chose to renew limited token usage at current levels, but the email warned that it is unclear “if the Army CIO pool will be renewed after 1 Oct.”

The Army Is Not Alone
This pattern is repeating everywhere. Meta, which encouraged employees to “tokenmaxx” and use AI tools aggressively, quietly removed its internal token usage leaderboard. Last week, Instagram CEO floated the idea of capping token use per engineer. Uber burned through a year’s worth of generative AI tokens in four months.
The common thread is simple: when you give humans a powerful, useful tool with no visible cost, they use it. Heavily. And the infrastructure behind “unlimited” is always finite.
For enterprise AI buyers, the lesson is stark. A sales pitch of “unlimited tokens” does not mean what it says. It means the provider believes you will not use enough tokens to exhaust the pool. If you do, the cap will be reinstated, the pricing will change, or the service quality will degrade.
What This Means for Heavy AI Users
If you spend $300 to $5,000 per month on AI API calls or subscriptions, you are already navigating the same dynamic at the consumer and prosumer level.
Claude Pro, Claude Max, ChatGPT Plus, ChatGPT Pro — every subscription tier has invisible limits baked in. Claude Fable 5 availability was capped at 50% of weekly usage during its redeployment period. ChatGPT removed its 5-hour usage limit for Plus and Pro users, but the Army story shows that removal is provisional. The underlying capacity is finite.
The practical reality is that every AI provider operates shared pools of compute. When you hit a usage threshold that makes your account unprofitable, the provider has three options:
- Reintroduce or tighten usage limits
- Raise prices
- Introduce tiered or metered pricing
Anthropic has already demonstrated all three. The June 15 billing change that would have moved Claude Code heavy users from subscription to full API pricing was only paused, not cancelled. Claude Sonnet 5 intro pricing ends August 31. These are not bugs. They are structural features of an industry where compute remains scarce at the frontier.
How to Protect Yourself
The Army is now rethinking its approach. You should too. Here are the practical steps for heavy AI users.
1. Assume Every “Unlimited” Plan Has a Ceiling
Build your usage model on the assumption that your current subscription tier has a hidden ceiling. When Claude Max costs $200 per month and the AI agent you run in the background consumes tokens at a rate that would cost $5,000 on the API, the subscription cannot last indefinitely at that price.
Track your token consumption yourself. Do not rely on the provider’s dashboard to tell you when you are approaching a threshold.
2. Diversify Across Providers
If you depend on a single provider for the majority of your AI workload, you are exposed to sudden policy changes. The Army had to negotiate with a single contractor. You do not need to.
The cheapest way to manage token exhaustion risk is to have a second provider ready to take over. Prompt caching differs across providers, but having a fallback is better than being forced into emergency negotiations.
3. Separate High-Value From Commodity Workloads
Not every AI task needs Claude Opus 4.8 or GPT-5.6 Sol. The cheapest models in the market are 95% cheaper than the most expensive ones with only 10-20% quality loss on many tasks.
Route your commodity workloads through cost-efficient providers and reserve frontier models for the work that genuinely needs them. Tools like OpenRouter, or simple model fallback logic in your own code, can cut your bill by 30-50% while maintaining resilience.
4. Watch for the Cue
The Army’s internal email warned about October 1. Your provider will give you similar cues if you pay attention. A change in how the provider talks about usage limits. A new FAQ entry about “fair use.” A pilot program for metered pricing.
The providers are signaling their constraints. Do not ignore the signals.
The Bottom Line
The Army’s token exhaustion is not a procurement error. It is the natural result of giving people a powerful tool without a feedback mechanism on cost. The same dynamic will play out in every large organization that adopts AI, and every heavy individual user who upgrades to a premium subscription and uses it in earnest.
Unlimited is a marketing term. Finite is the engineering reality. Plan accordingly.

Now available
Stop guessing your AI limits
The Mac app and web dashboard watch your Claude, ChatGPT, Gemini and more, and warn you before quotas hit.