9 min read B2B FinOps

AI Spending Needs Hard Budget Caps by Default

Searches for AI spending jumped 45% as AWS and Google Cloud finally ship spend limits. Why hard budget caps beat soft alerts for your agents.

AI Spending Needs Hard Budget Caps by Default

Searches for “ai spending” are up 45% in three months, and “spending limits” is up 115%. That is not abstract curiosity from finance departments. It tracks a specific fear that has moved from the enterprise boardroom to individual developers in 2026: an agent you launched, walked away from, and forgot about, quietly burning money while you sleep.

On October 3, Simon Willison published the sharpest articulation of the problem so far under the headline “We’re going to need default hard budget caps on pretty much everything.” His argument is simple and, for anyone running coding agents, uncomfortable. Soft caps do not work. A warning email at $80 is not a control. A hard cutoff is. The rest of this article is about what that means if your AI spending runs through Claude Code, ChatGPT, Codex, Cursor, or any hosted API you have wired into an agent loop.

What a hard budget cap actually is

A soft cap tells you that you have crossed a line. A hard cap refuses to let you cross it. The distinction sounds pedantic until you have woken up to an overage you did not authorize.

Willison’s framing is worth quoting almost directly: pay-by-usage services should let you say “after $X per month, cut this thing off and return errors.” If your usage hits the ceiling, the system pauses instead of continuing to bill. You get a failed request, not a surprise invoice. Errors, he argues, are preferable to a $10,000 bill nobody approved.

That matters more in 2026 than it did in 2024 for one structural reason. Coding agents and personal agents have collapsed the friction of spinning up code that spends money. A single prompt can now scaffold a service that calls paid APIs, provisions hosted storage, and autoscales compute. The thing that used to require a planning meeting now takes one reply. Traditional budget alerts, built for humans who deploy deliberately, were never designed for a loop that can spawn ten thousand billed requests before lunch.

AWS and Google Cloud finally shipped spend limits

The most interesting twist in this story is that the cloud providers people have been asking for exactly this feature for a decade have, in the last few months, started shipping it.

In July 2026, Google Cloud added spend caps to its budgets tooling. Budgets had long been able to send alerts and to flag early anomalies in cost data. The update let them enforce a ceiling rather than just report on one. Google’s own framing was aimed at AI spend specifically: the workloads that scale unpredictably and quietly.

AWS followed in September. Spend limits arrived for the new builder experience, described as project-level ceilings that pause your project when usage reaches the limit. The documentation is unusually candid about the tradeoff: spend limits are designed for experimentation, learning, and sandbox workloads, and AWS notes they can be used in production “when it is acceptable to have a brief pause of your resources.” Read that as an admission that a hard cap necessarily trades uptime for cost certainty, and that the company expects you to want both and to have to choose.

Two caveats matter before you assume you are covered. First, AWS’s spend limits are still rolling out to “a limited number of customers,” and the docs warn you may not have access yet. Second, the limit applies to a project, and your account can hold projects with limits and projects without them. A cap on one project does nothing for the metered API key you forgot lives in another.

Why “agent scope” is where AI spending really escapes

Here is the part the provider announcements do not solve for you. When a coding agent runs for hours, it does not spend in one predictable line item. It fans out. It calls a frontier model for planning, a cheap model for a classification step, an embedding endpoint for retrieval, and a hosted vector store that bills by the gigabyte-hour. A spend limit on your cloud project can catch the infrastructure half. It rarely catches the model API half, because that spend runs through a separate vendor with its own separate ceiling, or no ceiling at all.

This is the gap Willison is pointing at and the gap the last four months of TokenKarma coverage keeps circling. OpenAI reimposed a five-hour Codex cap in September. Anthropic layered weekly limits on top of rolling session allowances. Cursor published its own research in late September claiming a 7% reduction in token cost for users through harness changes alone. Every provider is now metering, capping, and re-metering, which tells you the cost pressure is real on both sides of the table.

The practical consequence for a heavy user is that “my AI spending” is not one number. It is a portfolio of meters, each with its own reset window, its own overflow behavior, and its own failure mode when you hit the wall. A hard cap on any single meter gives you a false sense of control while the others run unbounded.

How to build your own hard cap today

You cannot wait for every vendor to ship the feature. You can, however, enforce a hard ceiling yourself, and the work is smaller than it sounds.

The core move is to put a wrapper around spend rather than trusting each vendor’s dashboard. Track every billed call you make in one place, whether that is a gateway, a proxy, or a simple logging table in front of your provider clients. Then make the process that issues the calls check a running total against a monthly ceiling before it fires. When the total crosses the line, the wrapper returns an error and stops. This is the same shape as the AWS spend limit, just implemented one layer up, where it can see every provider at once.

Three specifics make it stick:

  • Set the ceiling per agent, not per account. A runaway loop is a runaway loop. If your research agent and your production agent share one bucket, the experiment can starve the revenue path. Separate buckets fail in isolation.
  • Make the default off, the override on. Willison’s strongest point is that unrestricted spending should be an opt-in someone ticks deliberately, not the silent default. Flip the polarity in your own tooling. New agents start capped. Removing the cap is a conscious act.
  • Alert at 70%, cut at 100%. The warning is still useful for the human, but it is not the control. The control is the cutoff. Treat the alert as a nudge to decide whether to raise the ceiling on purpose.

None of this replaces provider-side limits, and you should turn those on wherever they exist. It closes the gap the providers leave: the spend that crosses vendor boundaries, which is exactly the spend that grows fastest once agents get ambitious.

The bigger shift: erroring out is becoming a feature

There is a cultural change buried in all this. For years, “the service never goes down” was the selling point, and a budget limit that could pause a workload was treated as a risk to exploit, not a feature to want. That framing is inverting.

Once agents can spend faster than a human can react, the ability to be stopped becomes more valuable than the ability to keep running. A hard cap that returns errors is not a limitation. It is the difference between a bad afternoon and a bad quarter. The providers that make this the default, rather than a buried setting, will win the builders who have already been burned once.

That is the direction the market is moving, unevenly. AWS and Google Cloud shipped caps this summer. Model providers meter aggressively but still leave the enforcement mostly in your hands. The gap in between is where your exposure lives, and it is not going to close on its own.

Set your ceiling today. Make the unrestricted path the one that requires a deliberate click. Then sleep, knowing that the worst case is a failed request, not an invoice you will be explaining to someone for the next three months.