
TL;DR: Observe has no false-positive cost and prevents nothing. Warn is only meaningfully different from observe if the warning arrives before the effect. Block is the only mode that prevents anything, and it is survivable only when scoped to a class of effect you can evidence precisely. The sequence that works is observe, define a narrow class from the evidence, block that class, widen.
Most discussions of agent enforcement collapse into a preference: some teams want to block, some want to watch. Framed that way it is unresolvable, because both positions are defensible in the abstract.
Framed as a question of evidence, it resolves. Blocking requires knowing which class of effect to block, and that knowledge only exists after observing. The interesting question is not which mode is right but how you get from one to the next without either shipping a control nobody trusts or watching indefinitely.
Table of Contents
- What Is Runtime Enforcement?
- The Three Modes and What Each Costs
- When a Warning Is Just an Observation
- The Confidence Threshold for Blocking
- Failure Handling: Fail Open or Fail Closed
- A Rollout Ladder
- Metrics That Tell You Whether It Is Working
- Where Enforcement Cannot Reach
- Frequently Asked Questions
What Is Runtime Enforcement?
Runtime enforcement is the decision, made while an agent is executing, about whether to record an effect, surface it to a human, or prevent it. It sits downstream of the permission decision, which has already been made, and upstream of the effect, which has not yet happened or has not yet become durable.
Its unit is the effect rather than the request, which is what separates it from an approval prompt. An approval prompt asks a human whether an action should proceed, before it runs, based on the action's description. Enforcement acts on what the action is doing, based on observation. The two are complementary, and the second is what remains when the first has already said yes.
The Three Modes and What Each Costs
| Mode | Prevents | False-positive cost | Fails by |
|---|---|---|---|
| Observe | Nothing | None | Producing evidence nobody acts on |
| Warn | Nothing directly | Attention, and desensitization | Arriving after the effect |
| Block | The effect | Broken developer workflows | Being switched off after a bad week |
Observe is the only mode with no cost to the people doing the work, which makes it the correct starting point and a comfortable place to stay too long. Its characteristic failure is organizational rather than technical: a complete evidence trail that no process consumes is a compliance artifact. If nothing in your operating rhythm reads the trace, observation is not a control.
Warn trades attention for coverage. It works when volume is low enough that each warning is read, and stops working above that threshold, at which point it is worse than observation because it consumes attention and returns nothing. The rule of thumb worth adopting is that a warning stream nobody triages within a day should be a dashboard instead.
Block is the only mode that prevents anything. Its cost is paid by developers when it is wrong, and that cost compounds: a control that breaks a build twice in a week gets an exception, and the exception gets copied. Blocking survives on precision rather than coverage.
When a Warning Is Just an Observation
This distinction is worth making sharply because vendors blur it, sometimes without meaning to.
A warning that arrives before the effect is a control: the human can act. A warning that arrives after is a notification about something that already happened, which is observation with a louder interface.
Codex's documentation is unusually candid on exactly this point. Its safety monitoring "runs asynchronously and can pause a task if it detects potentially unsafe model behavior," and OpenAI states plainly that "a pause can arrive after the activity that triggered it" and that monitoring "doesn't replace sandboxing, permissions, or review of the result." Cursor is similarly direct that its Auto-review classifier "is not a security boundary" and can allow a call you would have blocked.
Both vendors are describing the same architectural reality: a judgment made in parallel with execution cannot reliably precede the effect. When you evaluate any warning-mode control, the question to ask is where in the effect's lifecycle the warning is emitted, and whether the effect is durable by then. If it is, classify the control as observation in your own documentation regardless of what the product calls it.
Runtime control infrastructure like Gensee Crate is designed around holding the effect at the boundary rather than reporting on it afterward, which is the architectural difference that makes the warn-versus-block distinction real rather than nominal.
The Confidence Threshold for Blocking
Blocking is worth doing when the evidence supporting the block is strong enough that a false positive is rare and explicable. Three properties make a class of effect blockable:
1. It is evidenced by a correlation, not a single event. "Credential file read" fires constantly during ordinary work and blocking it is unworkable. "Credential file read, then an outbound connection from the same process to a host approved in a different session" is rare, specific, and hard to produce accidentally. The first is an event; the second is a correlation, and only correlations support blocking.
2. It has a clean explanation when it fires. A developer who hits the block should be able to read one sentence and understand what happened. If explaining a block requires a paragraph about heuristics, the class is too broad and it will be treated as noise.
3. It has an escape hatch with a cost. A block with no override breaks legitimate work and gets removed entirely. A block with a frictionless override is not a block. The useful middle is an override that is available, logged, attributed, and reviewed.
Gensee's trace analysis of eight autonomous agent runs is a good illustration of where these thresholds come from in practice: what separated the boundary-crossing runs was service semantics, independent evidence and persistence rather than activity volume. Those are correlation properties, and they are the shape of a blockable class. A rule built on volume would have fired on the safe runs too.
Failure Handling: Fail Open or Fail Closed
Every enforcement layer eventually cannot run, and what happens then is a design decision that deserves explicit choice rather than inheritance from a default.
The agent vendors have made different calls here and they are instructive:
- Claude Code's Bash sandbox fails open. If it cannot start because dependencies are missing or the platform is unsupported, Claude Code warns and runs commands unsandboxed.
sandbox.failIfUnavailable: trueconverts that to a hard failure, and Anthropic's documentation notes this is intended for managed deployments that require sandboxing as a security gate. - Cursor's Linux sandbox fails closed to a prompt. If the kernel lacks Landlock v3 support, Cursor falls back to asking for approval before running commands.
Neither is wrong. Failing open preserves developer velocity and silently removes the control; failing closed preserves the control and produces friction that a developer experiences as the tool being broken. What is wrong is not knowing which one your deployment does.
The guidance that holds up: fail open during the observe phase, where the control provides no protection anyway and an outage should not stop work. Fail closed once you are blocking, because a control that disappears under exactly the conditions an attacker might induce is not a control. And make the transition an explicit, dated decision rather than a config default.
A Rollout Ladder
Six steps, in an order where each produces the evidence the next one needs.
1. Observe everything, act on nothing. Run across a real workload for long enough to see normal variation, including release weeks and incident weeks. You are building the baseline that defines what "unusual" means.
2. Find the correlations, not the events. Work through the trace looking for effect sequences rather than individual actions. The output of this step is a shortlist of candidate classes, each described as a correlation.
3. Warn on one class, to a channel someone owns. One class, not five. If the warning volume exceeds what that owner triages, the class is too broad; narrow it and repeat rather than proceeding.
4. Measure the false-positive rate before blocking anything. If warnings on the class fire on legitimate work more than rarely, it is not ready. This is the step teams skip, and it is the one that determines whether the block survives.
5. Block the class, with a logged override. Announce it, document the one-sentence explanation a developer will see, and make the override attributable and reviewed.
6. Widen one class at a time, repeating steps 2 through 5. Never widen two at once, because when something breaks you will not know which.
Gensee's 15 questions to answer before deploying AI coding agents is the broader readiness checklist this ladder assumes has been worked through.
Metrics That Tell You Whether It Is Working
Four numbers, and one to avoid.
- Override rate per blocked class. Rising means the class is too broad. Near zero over a long period means it may be too narrow to be worth the operational cost.
- Time from warning to triage. The single best predictor of whether warn mode is real. Above a day, it is a dashboard.
- Coverage gaps observed. How often an effect was recorded through a channel your enforcement layer does not mediate. This number should be tracked deliberately, because it is the honest measure of what the control does not do.
- Exception age. Exceptions granted to unblock work should have expiry dates. The distribution of their ages tells you whether the control is being managed or eroded.
The metric to avoid is total events blocked. It is the number most likely to be reported upward and least likely to mean anything, because it rises with both better detection and worse precision, and there is no way to tell which from the number alone.
Where Enforcement Cannot Reach
Three limits belong in your own documentation, stated before someone else finds them:
- Enforcement acts on mediated effects. Anything reaching the outside world through a channel the layer does not intercept is unenforceable, and knowing precisely where that line falls is the most important technical property of any implementation.
- Enforcement acts on effects, not intent. It can hold a write and record a read. It cannot determine whether the developer meant it, and treating a block as a finding about a person rather than about a sequence is a fast way to lose the team's cooperation.
- Enforcement within a session cannot see across sessions unless something carries lineage. A block evaluated on this session's evidence alone will miss a sequence begun last week, which is the failure mode Gensee's multi-session agent safety research documents.
Frequently Asked Questions
Should we start by blocking or observing?
Observe. Blocking requires knowing which class of effect to block, and that knowledge only exists after observing a real workload. A block defined before the baseline is a guess, and a guess that generates false positives gets switched off permanently.
What makes an effect safe to block?
Three properties: it is evidenced by a correlation rather than a single event, it has a one-sentence explanation a developer can act on, and it has an override that is available but logged and reviewed. Classes missing any of the three are not ready.
Is a warning mode actually useful?
Only if the warning arrives before the effect becomes durable and someone triages it within a day. Both Cursor and OpenAI document that their classifier and monitoring judgments can arrive after the fact; where that is true, the control is observation with a louder interface, and should be documented as such.
Should the enforcement layer fail open or fail closed?
Fail open while observing, since the control provides no protection anyway and an outage should not block work. Fail closed once you are blocking, because a control that vanishes under inducible conditions is not a control. Make the switch an explicit decision.
What metric should we report on?
Override rate per class, warning-to-triage time, observed coverage gaps and exception age. Avoid reporting total events blocked; it rises with both improved detection and degraded precision, and cannot distinguish them.
Conclusion
The mode is not a philosophy, it is a function of evidence. Observe until you can describe a class of effect as a correlation with a clean explanation, block that one class with a logged override, then widen deliberately. Decide the failure behavior explicitly, and track the coverage gaps as a first-class number rather than an embarrassment.
Gensee Crate is built to hold effects at the boundary rather than report on them afterward, and to carry the lineage that makes a cross-session correlation expressible in the first place. To discuss a deployment, book a demo or see pricing.