How do you stop every AI agent at once?
covey has two switches: one per agent and one per organisation. Both set a flag, abort the running session and put the started task back into the backlog.
Several agents are working, and something has just gone wrong with one of them. Without a switch for it you would have to find every running session, find the container behind it, tear that down, and set the started task back to open by hand. Then miss the one that wakes shortly afterwards out of its own schedule.
Such a case is on record. On 8 May 2026 a user posted the log of their coding agent to Hacker News, item 48069567. The agent had run rm -rf ./* in the wrong working directory and deleted every local git repository with it (a Timeshift snapshot undid most of the damage, at a cost of one to two hours). The agent's message to the team sits in the same log: "I will stop all further actions immediately."
That is the good case. The agent noticed what it had done and stopped itself. Just under four months earlier another user on Hacker News had asked about the bad one, item 46601809: whether IAM policies, approval workflows and monitoring are enough or whether there is a gap. Covey's documentation carries the same question as the one that comes up in every security review: "And if we have to stop all of this?"
The switch on the agent sets a flag and aborts the session
POST /api/v1/agents/{id}/kill works through a fixed order. The agent's killed flag goes true in the database. A message of type kill with the reason kill-switch reaches the daemon in the sandbox. The running session is cancelled. A parked warm sandbox is evicted on the spot by evictWarm, because a stopped agent must not keep a container running.
The step after that is the one where work could be lost. An UPDATE backlog_tasks SET state='open' pulls every task of that agent out of in_progress and back into the backlog, so the aborted run ends and the started task stays in the backlog. A lifecycle entry with status killed goes into the recording afterwards.
On the far side of the connection this is rough. The daemon cancels the run under way and returns ErrKilled, and coveyd then writes the line "kill switch" with the addition "daemon terminates immediately" to the log and exits with code 3 (code 1 is the ordinary error path).
The second switch sits on the organisation
POST /api/v1/fleet/kill first sets a single column: UPDATE organizations SET fleet_killed=$2 WHERE id=$1. After that it calls the same kill for every agent of the organisation. Who may do that is drawn tightly (org_admin or security). Only that role sees the red button in the dashboard, and whether it has been pressed is answered by GET /api/v1/fleet in the field fleet_killed. That one column already holds the whole workforce. Before an agent is woken the control plane reads its own flag first and the organisation's second; if either stands, it returns without starting a sandbox.
Both flags sit in the query the scheduler collects its due heartbeats through: WHERE NOT a.killed AND NOT org.fleet_killed. Whoever fires a heartbeat by hand in the interface instead gets a 409 back, with the sentence "Agent or fleet is stopped" and the instruction "resume it first."
A trigger from outside via POST /api/trigger/{token} bounces as well, though there it bounces off the agent's own flag, which KillFleet has set on every agent individually.
TestNotausDerFlotte in internal/integration/notaus_test.go checks that a new task no longer wakes the stopped agent. A task is created while the emergency stop stands, and two seconds later it is still in state open. A comment above the test gives the reason it was written. The fleet-wide stop had no test at all, "of all things the lever you need exactly once and that has to work then".
Releasing has to reach as far as stopping
A mistake sat here, and it stands today as a comment above ResumeFleet in internal/orchestrator/orchestrator.go. Taking the stop back at first released only the organisation's flag. KillFleet had stopped every agent individually, though, and those flags stayed up. The interface then reported "no emergency stop" while the workforce stood still.
Today ResumeFleet clears both the organisation's flag and the flag of every single agent. Whoever has open work is woken right away instead of on the next tick.
In the other direction the rule works differently. Releasing a single agent while the fleet switch stands does reset that agent's own flag; the check against the organisation's flag still sits in front of it.
What the emergency stop costs
Price and promise sit in the same sentence on the guard-rails page: both switches take effect immediately, and running sandboxes are torn down. What lies in the sandbox home is therefore the state of the last sync. According to spec/16-runner.md that sync runs at real falling-asleep, and a parked warm sandbox is secured in the background at most once every five minutes.
The fleet switch is coarse on purpose, and that is the second price. KillFleet collects its agents with SELECT id FROM agents WHERE org_id=$1, so every one that belongs to the organisation.
Systemic trouble is what the tool is built against (an injection, say, that has hit more than a single agent). Whoever suspects one agent takes the switch on the agent.
Point three is about who hears of it. Pulling the fleet switch emits a notification of class cost and kind fleet_killed, "The emergency stop was pulled: every agent of this organisation is paused", linking to /administration. The recipients according to spec/06-observability-control.md are the org admin and controlling.
That agent in the 8 May log stopped itself and then asked for help. A switch outside the runtime answers the other case: the agent that stays silent.