Why don't an agent's limits live in its prompt?
A rule in SOUL.md binds the agent through the prompt. A guard rail sits at the broker, the action layer and the approval queue, fail-closed.
An agent needs a limit, and the obvious place to put one is its configuration. SOUL.md has a section for exactly that. Under ## Limits you then find a line like this one: "No actions that delete customer data without approval."
The line is not useless — it sits in the system prompt of every run and steers behaviour from there.
Then a ticket arrives whose text reads like an instruction. Or a mail somebody prepared. The rule is still there, but it has become one sentence among other sentences.
A rule in the prompt binds only those willing to follow it
Self-binding is what the spec calls limits like these. They steer behaviour and they are worth having — a security boundary they are not, because a prompt can be worked around or defeated by injection.
That is not a fault of the text but of who interprets it. What an injection compromises is the agent's own reasoning, and that same reasoning decides what the rule is worth.
Hard limits therefore live outside the runtime. The agent cannot read them, cannot argue them round and cannot bypass them. All it notices is that an action was refused.
The prompt does not disappear over this. Taken together the two layers are what the spec calls defence in depth.
Seven places the platform already sits in
A guard rail needs no new control point. It takes effect where the control plane and the daemon already carry the data flow, and that comes to seven places.
The secrets broker decides which system and which scope an agent can get a token for. Egress decides who the sandbox is allowed to reach. Inside the daemon, the tool and action layer decides which action runs.
Four more follow: the approval queue, rate and cost limits, a content filter for PII, and the style gate.
Rules apply on three levels: globally for every agent, for a role, or for a single agent. Additive-restrictive is the principle behind that. A narrower level tightens a rule; it never softens a global deny.
The security and compliance role administers them, deliberately separate from the people who own the agents. Otherwise a single team lead could soften the org-wide limits.
In doubt it is refused
The default is fail-closed. What is not allowed is forbidden, and in doubt an action is blocked or sent for approval.
After installation the defaults show the pattern. An outbound reply to a customer needs an approval. HR systems are off limits. Anything called delete is hard-denied.
An approval is not an abort. Through request_approval the daemon reports it, the control plane halts the action, and the agent sits in blocked meanwhile. Waiting costs it no compute, and it carries on as soon as the decision arrives.
Budget per agent follows the same logic against a different mistake. Cumulative cost is what it measures, which makes it a lifetime ceiling instead of an allowance per run. Once it is passed the agent is paused and the running task goes back into the backlog. We have not built a cap per run or per time window.
None of this is the agent's to change, because the rules live in the control plane and not in its configuration. A human with the right role changes them, and that change lands in the audit trail afterwards.
The style gate measures what goes out
One guard-rail type is called style_gate, and it shows that the mechanism is not only about security.
A TONE.md describes an agent's voice. Like every prompt rule it is self-binding: nothing checks whether the text leaving the sandbox follows it. Checking precisely that is the style gate's job, and it measures the free text of an action before the action runs — a GitLab comment, a mail.
The yardstick is the profile block in TONE.md, schema covey-style/1: bands per metric plus absolute floors such as a sentence over 35 words. Only a HIGH finding makes the gate act: a metric a full band width outside its band, or over an absolute floor. Text under 60 words it does not measure. A one-line comment has no style.
Under deny the action is refused and the findings are the reason: which paragraph, which metric, which band. Revising and trying again is then the agent's move. After two denials on the same task and action the text goes to the approval gate. A loop that does not converge therefore ends with a human and not at the turn limit.
Two of these limits we drew deliberately. First, the gate is a measurement and not a security boundary: with no profile configured it records that it did not apply and lets the action pass. Second, it sits only on the allow path, and it never softens a refusal that came from another rule.
So much for the price we pay. A gate that measures instead of judging sometimes refuses a text that was fine.
What remains at the end is more than a prevented action. An attempt that fails at a guard rail lands in the recording afterwards as a refused attempt, together with the arguments of the call. A rule in the prompt leaves no such event. The action that broke it sits in the recording with its arguments, and nothing there marks it as a breach. Put the rule in the control plane and every trigger writes one.