How do you run Claude Code on a server without sitting in front of it?
Claude Code goes headless with one flag. What a server run needs beyond it: a key outside the container, a session that waits, and limits on the call itself.
An agent that works at night needs a machine that runs at night. The usual route is well known: rent a small box from a hoster, install Claude Code, clone the repository.
A private network comes on top so no port stands open, and tmux so the session survives a dropped connection. Then you can close the lid and let the box work the night alone. It works.
What is missing shows up in the first weeks of operation. On a Linux server the token sits in ~/.claude/.credentials.json with file mode 0600. That mode keeps other accounts out, and every process under the same account (the cron job, the deploy hook) reads the file anyway. Nobody has settled how many steps a run may take.
What happened in that window is afterwards held in the scrollback buffer alone, and tmux keeps 2000 lines by default. And when the customer replies in the morning, nobody is sitting there to type that reply into the waiting window.
That gap is where covey starts. We use the same mode — the difference is the operation around it.
Headless is a flag, the operation is everything else
Claude Code has a non-interactive mode: claude -p "…" runs a job without an interface. Inside the sandbox image a daemon drives the tool through an adapter, headless, with --output-format stream-json.
More than the result comes out of that stream. Every line is a JSON event (assistant, tool_use or result) that the daemon forwards to the control plane. Recording a run therefore costs almost nothing — whoever looks later sees every tool call instead of a summary.
The step limit hangs off the same call. --max-turns carries it, and without one of its own the orchestrator substitutes thirty. The budget travels along on the run request as a MaxBudgetUSD field, which the adapter leaves lying when it builds the arguments.
The cost limit works one level up. The closing result event names total_cost_usd, which the daemon reports to the control plane. There the orchestrator holds the agent's running total against its BudgetUSD field and the budget_limit guard rule, whichever is stricter, and once that is reached it pauses the agent, reopens the task with "budget exceeded" and sends the daemon a TypeKill with the reason budget.
The check waits on that report, which Claude Code delivers at the end of a run, so between two reports the total can overshoot the limit.
On a subscription seat that figure is notional, because Claude Code prices the run as if it had been billed while the seat is paid anyway. The platform books it unchanged and labels it with the credential it came from.
The key does not live in the sandbox
Nobody has ever logged in interactively in a freshly started sandbox. With no secret deposited there are no credentials in it, and every task fails on "Not logged in · Please run /login", which the adapter rewrites into a pointer at the missing secret. Three forms are provided for: an API key in ANTHROPIC_API_KEY, a long-lived OAuth token (generated once with claude setup-token), and provider credentials for Bedrock, Vertex or Foundry.
Which name fills which variable is settled rather than guessed from the token's prefix: anthropic_api_key becomes ANTHROPIC_API_KEY, claude_code_oauth_token becomes CLAUDE_CODE_OAUTH_TOKEN.
What separates this from the box under the desk lies elsewhere: the key is a brokered secret, injected into the daemon at runtime rather than sitting permanently in the sandbox.
A session that waits for tomorrow's answer
Headless runs are stateless to begin with, but they can be threaded. The call returns a session_id, and a later run with --resume <session_id> loads the context of the previous one.
Covey uses that for the case a tmux window cannot cover.
When the agent asks a follow-up question, the daemon reports the task as blocked and stores the correlation key with the session_id on the parked task. When the ticket reply arrives, the sandbox starts again and the adapter calls claude -p --resume <session_id> with it. The agent is not waiting, it is asleep, and asleep it costs nothing but a row in the database.
The limit of that session is stated in spec/12 beside it. It is short-term working context, a bounded context window that can expire. Durable knowledge lives in the platform's memory layer, in the agent's wiki pages.
What survives the container, and what does not
We create the container for the wake and throw it away afterwards. What stays is /home/agent: cloned repositories, downloaded attachments, the wiki pages and ~/.claude. The run starts with that directory as HOME.
That gives a rule for every tool an agent needs: toolchain into the image, version into the home. The home is walked at every wake and written back after every run; the image is pulled once per runner. The dev-flutter image reverses that rule on purpose: in dev, fvm fetches the ~1.3 GB SDK into every Flutter agent's home, while the role image carries the baseline version, because for a Flutter agent the version question is settled and fvm stays installed beside it.
The price of this arrangement is concurrency. An agent handles one task at a time, because two concurrent runs of the same agent would fight over the same home. Throughput comes from more agents running beside each other. Saving startup time takes a warm sandbox: the container stays up between two wakes, costs memory and saves seconds.
You can close the lid in either arrangement. The difference shows the morning after — in whether anybody can read back what happened overnight, and whether the run carries on by itself where the customer's reply waits.