Jahanzaib
Back to Blog
Trends & Insightsai newsai-agentsai-costs

The Army Promised Unlimited AI Tokens. It Ran Dry in Weeks.

The Army said its AI tokens were unlimited, then hit the wall in about five weeks. Here is what really drained the pool, and the boring budgeting that keeps your own AI bill from doing the same thing.

Jahanzaib Ahmed

Jahanzaib Ahmed

July 23, 2026·11 min read
The Army Promised Unlimited AI Tokens. It Ran Dry in Weeks.

In May 2026 the US Army told its people the AI faucet was open. Unlimited tokens. Use as much as you want, and please, use more. By mid-June the pool was dry and the limits were back. That's not a years-long budget cycle. That's about five weeks.

I've been building AI systems for businesses long enough to recognize the shape of this one on sight, because it's the same wall my client hit last quarter, just with more zeros. The first time this really bit me, it was a client running a customer support bot on a generous monthly plan who drained the whole month's budget in nine days, most of it spent on an expensive reasoning model answering questions a cheaper model would have nailed. In my experience the meter always moves faster than the demo promised.

The Army just proved it at a scale most of us will never touch. It had one hundred million tokens for the year, and it burned the lot on a single service before summer.

Wired and Ars Technica both broke the details this week, and the story is funnier and more useful than the "government waste" headline suggests. The word doing all the damage here is "unlimited." It's a marketing word. It is never an architecture. If you're planning to put AI in front of your team, this is the cautionary tale to read before you sign anything.

Key Takeaways

  • The US Army offered "unlimited" AI tokens in May 2026. By mid-June its enterprise token pool was exhausted and per-user limits came back.
  • The Army had 100 million tokens for the year across Ask Sage, its enterprise AI workspace. One anonymous employee said the whole Army "burned through the whole year of tokens for just one service."
  • Users got at least 200,000 tokens a month and were automatically topped up when they ran out. People who signed up but didn't use it were nudged to burn more.
  • This isn't a military problem. Meta pulled its internal "tokenmaxx" leaderboard and is capping use, and Uber engineers reportedly spent a year's token budget in four months.
  • The fix is boring and it works: estimate tokens per task, budget per user, pick the right-sized model for each job, and monitor consumption instead of trusting an "unlimited" label.
Ask Sage homepage, the US Army's enterprise generative AI workspace, showing model support and a token usage monitor in the interface
Ask Sage, the platform at the center of this. Note the "Monitoring Token Usage" item in its own sidebar. The tooling saw it coming.

What actually happened with the Army's AI tokens?

The Army CIO announced unlimited tokens in May 2026 and had to walk it back within weeks. According to an email viewed by Wired, "by mid-June the Army CIO pool was exhausted of tokens and had to re-establish limits." The Army chose to renew usage at current levels, but even that came with a caveat: it's unclear whether the pool gets renewed after October 1. So the "unlimited" era lasted roughly a month before reality showed up with an invoice.

The platform is Ask Sage, a generative AI workspace accredited for Controlled Unclassified Information and used across the Department of Defense. It runs a menu of models including Google's Gemini, Meta's Llama, and OpenAI's ChatGPT. The Army bought into it with an annual "enterprise pack" of 100 million tokens. An Army employee, speaking anonymously to Wired, put it plainly: "Apparently the whole Army burned through the whole year of tokens for just one service." Each user got at least 200,000 tokens a month with automatic top-ups, and idle sign-ups were actively pushed to use more.

The Army's official line, from Colonel Marty Meiners, leans into it: "The transition from initial testing to widespread operational use proves the success and demand for these tools." Maybe. Or it proves that when you tell three and a half million people to "tokenmaxx" and remove the price signal, they will happily drain any pool you give them, useful work or not. One of the employees Wired spoke to said they hadn't found the tools particularly useful for their actual job. Both things can be true at once, which is exactly what makes this a budgeting story and not a technology story.

Wired article headlined The Army Is Burning Through Its AI Tokens with illustration of coins on fire
Wired's framing said it best. Coins, on fire, held up by tiny soldiers. The token meter does not care about your enthusiasm.

How does an entire Army burn through 100 million tokens so fast?

Because 100 million tokens is a puddle, not an ocean. It sounds enormous until you do the arithmetic. A token is roughly three-quarters of a word. One hundred million tokens is around 75 million words. Spread across an organization where hundreds of thousands of people are being told to use AI daily, that's a rounding error, and it disappears in weeks.

Here's the math that trips everyone up. A single involved chat, with a long document pasted in and a few back-and-forth turns, can easily run 20,000 to 50,000 tokens once you count the input, the context, and the output. Give one user a 200,000 token monthly allotment and they hit it in a handful of real work sessions. Now multiply by tens of thousands of users. The pool doesn't stand a chance. If you want to see how the numbers move for your own setup, our AI agent cost calculator lets you plug in calls per day and watch the monthly total climb faster than intuition says it should.

Two things make it worse than the naive estimate. First, reasoning models. The current generation "thinks" before it answers, and that hidden reasoning is billed tokens you never see in the reply. A question that used to cost 500 tokens can cost 10,000. Second, agents. An agentic loop that reads a document, plans, calls a tool, reads the result, and tries again can chew through hundreds of thousands of tokens on one task, because every step re-sends the growing context. If the Army had autonomous agents in the mix rather than plain chat, the burn rate makes even more sense. This is the same dynamic I walk through in our guide to agentic AI for business: autonomy is powerful and it is expensive, and the two arrive together.

Ars Technica article headlined Unlimited AI tokens aren't unlimited after all as US Army burns through supply
Ars Technica's headline is the whole lesson in nine words. "Unlimited" is a pricing decision someone else can reverse.

Why is "unlimited" AI pricing a trap?

Because unlimited removes the one signal that makes people budget: the price. When every query feels free, nobody weighs whether they actually need the expensive reasoning model to reformat a paragraph. Usage isn't driven by value anymore, it's driven by convenience, and convenience has no ceiling. The provider eventually notices the pool draining and pulls the plan, and now you've built workflows on an allowance that just vanished.

The Army is in good company here, which is the part that should get your attention. Meta ran an internal culture of "tokenmaxx," encouraging staff to rack up AI usage, then quietly took down the leaderboard tracking it and started curbing use. Adam Mosseri, who runs Instagram, floated capping token use per engineer. Uber, per Fortune, watched its engineers burn through a year's worth of AI tokens in four months. These are some of the most sophisticated technology organizations on the planet, and they all made the same mistake: they treated a metered resource as if it were free because a plan told them it was.

The uncomfortable truth is that "unlimited" plans are a bet the provider makes on your restraint. When a whole cohort of customers stops being restrained, the plan changes. Building on top of it without your own budgeting is like building on rented land and never reading the lease.

So don't outsource your budget to a pricing page. Own the number yourself.

Unlimited mindset versus budgeted token economics

The difference between the organizations that get burned and the ones that don't isn't the size of their AI bill. It's whether they treat tokens like a resource with a meter. Here's the contrast, drawn straight from this week's news.

DecisionUnlimited mindsetBudgeted token economics
Access model"Use as much as you want"Per-user and per-team budgets, visible to users
Model choiceBest model for everythingCheap model for routine work, premium only when it earns it
Reasoning and agentsOn by default, unmeteredReserved for tasks that justify the token multiplier
VisibilityNobody sees consumption until the pool diesLive dashboards, per-workflow token tracking
PlanningEstimate: "it's unlimited"Estimate tokens per task, model the monthly total up front

None of the right column is exotic. It's the same discipline you'd apply to cloud compute or any other consumption-priced resource. The mistake is treating AI as a flat subscription when it behaves like electricity.

Jahanzaib AI Agent Cost Calculator showing use case selection and monthly cost estimation for production AI agents
Pick a use case, set your call volume, and the monthly total appears. Doing this before you deploy is the whole ballgame.

How do you budget AI token costs before they blow up?

Start by estimating tokens per task, not per user. Pick your three or four most common AI jobs, run each one a few times, and record the actual token count including reasoning. That single exercise turns "unlimited" into a real number, and the real number is usually sobering. Once you have per-task costs, multiply by realistic volume and you have a monthly estimate you can defend.

From there, a few moves keep the bill sane. Match the model to the job, because using a frontier reasoning model to summarize an email is like renting a crane to hang a picture. Cap and monitor per user, so one power user's experiment doesn't drain the shared pool. And be deliberate about where you deploy full agents versus simpler automation, because the token cost gap between them is enormous. We break down exactly that tradeoff in when to use AI agents versus automation, and for teams comparing build routes, n8n agent workflows give you a place to see and cap consumption per step. The underlying concepts, from tokenization to the context window that quietly inflates every agent call, are worth understanding before you commit a budget to them. If your workloads lean on retrieval, our RAG guide covers how context size drives cost there too.

What does this mean for your business?

It means the question isn't "can we get unlimited AI," it's "do we understand what we're actually going to consume." The Army had a bigger budget than almost any company will ever have and it still hit the wall in five weeks, because scale plus enthusiasm plus no price signal equals empty pool. Your numbers are smaller, but the equation is identical.

That equation is not a reason to panic. It's a reason to plan.

The good news is this is a solvable planning problem, not a reason to sit out AI. Model your consumption before you deploy, right-size your models, cap per user, and watch the meter. If you want a running start, plug your real use case into our AI agent cost calculator to get a monthly estimate and a payback period in about a minute, and if you're trying to figure out where AI fits in your operation at all, the AI readiness assessment walks you through it. The Army's mistake was treating "unlimited" as a plan. Don't. Treat it as a warning.

How many tokens did the US Army actually have?

The Army had 100 million tokens for the year as part of an annual enterprise pack on Ask Sage. It announced unlimited access in May 2026 and exhausted the pool by mid-June, then reinstated per-user limits. Each user received at least 200,000 tokens a month with automatic top-ups.

Why did the tokens run out so quickly?

100 million tokens is roughly 75 million words, which is small once tens of thousands of people use AI daily. Reasoning models and autonomous agents make it worse: hidden reasoning and repeated context re-sends can turn a 500-token question into tens of thousands of tokens per task.

Is this just a government problem?

No. Meta encouraged internal "tokenmaxx" usage, then removed its tracking leaderboard and moved to cap consumption. Uber engineers reportedly spent a year's token budget in four months. The pattern shows up anywhere an organization treats a metered resource as if it were free.

How do I estimate my own AI token costs?

Run your most common AI tasks a few times and record the actual token counts, including reasoning tokens. Multiply by realistic monthly volume to get a defensible estimate. A cost calculator that accounts for model choice, call volume, and infrastructure gives you a fuller total cost of ownership.

Should "unlimited" AI plans make me nervous?

Treat them with caution. Unlimited removes the price signal that keeps usage tied to value, and providers reserve the right to reinstate limits once a cohort drains the pool. Build your own per-user budgets and monitoring so a plan change doesn't break your workflows.

Citation Capsule: The US Army announced unlimited AI tokens in May 2026 and exhausted its 100 million token annual Ask Sage pool by mid-June, reinstating per-user limits; users received at least 200,000 tokens monthly. Meta and Uber hit similar overruns. Wired (Jul 21 2026) · Ars Technica (Jul 22 2026) · Fortune on Uber token spend (May 2026) · TechCrunch on Meta token caps (Jul 2026).
Feed to Claude or ChatGPT
Jahanzaib Ahmed

Jahanzaib Ahmed

AI Systems Engineer & Founder

AI Systems Engineer with 126 production systems shipped. I run AgenticMode AI (AI agents, RAG systems, voice AI) and ECOM PANDA (ecommerce agency, 4+ years). I build AI that works in the real world for businesses across home services, healthcare, ecommerce, SaaS, and real estate.