Skip to content
covey

In February 2026 someone wrote on Hacker News that they had left an agent running before going to bed. It got stuck in a loop. By morning it had burned $200 in model calls. Eight weeks later another poster asked whether any tool stops LLM calls at runtime instead of only monitoring them, because most of what they had found reported the spend after it had happened.

Anyone running more than one agent needs a number the platform checks itself, and a consequence that happens while nobody watches. In covey that number is budget_usd. It sits in the agent registry beside the model and the turn limit, and a human with the right role sets it through POST /api/v1/agents/{id}/budget. The documentation lists it beside max_turns and the guard rails as one of the limits that sit outside the model and cannot be argued away.

The figure is made at the end of a run

The daemon in the sandbox sends a cost message to the control plane after every run. That message carries the amount in dollars, the token counts and the model name, and it goes out only when something was measured at all (cost or tokens above zero). The control plane books the entry and checks the cap in the same step.

The input side is counted in three parts: fresh input, a read out of the prompt cache, and writing the cache. With a runtime that caches its context the uncached share is the smallest of the three. One agent measured in the spec came to 5,497 input tokens against 1,842,222 output tokens, while a single one of its runs read 2.34 million cached tokens. A cache read is priced at about a tenth of fresh input, so the three are stored apart and added up only for display.

What comes out is a list-price equivalent, and the spec labels it as one. On a metered API key that is what was billed. On a subscription seat it is what the same work would have cost through the API, while the seat is paid for anyway. The number overstates the spend there, and it overstates it in the harmless direction: agents look more expensive than they are. The provider's own billing view stays the record.

Two places set the cap and the smaller one wins

The amount lives in two places. One is the field on the agent. The other is a guard rail of type budget_limit that carries its value under params.usd and can apply to a single agent or to everyone.

Whichever of the two is smaller is the one that binds. With nothing on the agent the guard rail applies on its own, and among several guard rails the smallest wins again.

It is measured against the agent's cumulative cost. The query sums every cost entry of that agent with no time window, which is what the spec means by a "lifetime ceiling". An allowance per run or per time window is something covey does not have today.

Knowing in advance where an agent will be stopped means looking in both places. The rule tester takes a system or a system:action, executes nothing, and returns the decision (allow, deny or require_approval), the rule that triggered and the smallest cap among the guard rails. It does not read the agent field, which is shown on the agent's own page instead. Asked about an agent whose field says 50 USD while a guard rail says 200 USD, the tester reports 200 USD, and the control plane still pauses the agent at 50.

Reaching it pauses the agent and returns the task

When the cap is reached the control plane records a guard-rail event with the rule, the limit, the amount spent and the consequence. Then it flips the agent's kill switch, sends the daemon a kill message and reopens the running task, which carries "budget exceeded — agent paused" as its reason.

The task is open in the backlog again. The source separates this outcome from a failure: the work stands, the agent is paused, and somebody has to decide what the number should be.

A notification of class cost goes out at the same time. Its title names the agent, reads "was paused", and carries the amount spent against the amount allowed; the link goes to /costs.

Getting back to work takes a human. POST /api/v1/agents/{id}/resume lifts the pause, under the same roles that operate the kill switch. If the cap stays where it is, the next cost entry pauses the agent again.

The cap acts between two runs

This is where the precision ends. A run's cost reaches the control plane once the run has finished, and the documentation says so plainly: the budget caps reactively, a runaway task can blow through it, and only the next one is throttled.

A sub-run in the project checkout books its cost over the same message as the outer run, so it cannot slip past the cap. Inside a single run the bound is the turn limit, --max-turns, 30 steps by default.

A budget flag for the run itself is listed in the spec, and the field travels as far as the daemon. The Claude Code adapter builds its argument list without that flag.

For the night in the example that means two things. The agent finishes its run, and what the run cost is known only afterwards. At the next wake it stands still, its task is open, and its own page at /agents/{id} shows the amount spent beside the cap set on the agent itself.

Back to all posts