← Back to all posts

Education & Learning

AI Agent Runtime Security: The Systems Guide

What runtime security means, how it differs from gateways and sandboxes, and the questions to ask before you buy one

Gensee Crate Team · · 12 min read

Security engineer reviewing a branching timeline of an AI coding agent session for runtime security risks

TL;DR: Runtime security is the layer that observes and governs what an agent actually did, as it does it, rather than what it was asked or allowed to do. It is a distinct layer from prompt filtering, AI gateways and OS sandboxes, and its defining property is that it spans sessions. If a control cannot connect an action today to one taken last week, it is doing something else.

The term "AI agent runtime security" is now attached to products that do very different things: prompt filters, API gateways, container sandboxes, observability tools and policy engines. That vagueness is expensive, because the layers fail differently and buying two of the same one leaves a real gap uncovered.

This guide defines the category narrowly, places it against the adjacent layers, and gives the questions that actually separate one implementation from another. It does not restate any particular product's architecture; where that is the subject, it links to the primary source.

Table of Contents

What Is AI Agent Runtime Security?

AI agent runtime security is the discipline of observing and governing an agent's actual effects at the moment they occur: the processes it starts, the files it reads and writes, the network destinations it reaches, the credentials it touches, and the changes it leaves behind. It operates on effects rather than on inputs or intentions, and it is defined by three properties:

  1. It is positioned after the decision. A permission system asks whether an action should be allowed. A runtime layer observes what the allowed action did.
  2. It is independent of the agent. Evidence produced by the process being observed is weaker than evidence produced beside it. This is the difference between a log and a record.
  3. It spans sessions. An agent's risk accumulates across many sessions; a control that resets at the session boundary cannot see the shape of it.

The third property is the one that separates this category from everything adjacent. Prompt filters, gateways and sandboxes are all session-scoped by design. If a tool cannot connect an action taken today to one taken last week, it may be a useful control, but it is not doing runtime security in this sense.

The Layer Map: What Runtime Security Is Not

Gensee's AI agent safety ecosystem map lays out the full stack. The short version, for the four layers most often confused with each other:

Layer Operates on Sees Blind to
Prompt filtering / model safeguards Text going in and out Instructions and responses What the tool call did after the text was approved
AI gateway API traffic to the model Requests, tokens, cost, model routing Local file reads, shell commands, anything not routed through it
OS sandbox One process's syscalls What this command may touch Why it touched it, what happened last session, effects outside its scope
Runtime security Effects, across sessions What actually changed and under whose authority Intent; it observes rather than reads minds

The gateway confusion is the most consequential in practice, because gateways market themselves in similar language. A gateway sits on the path between the agent and the model. Most of what a coding agent does never crosses that path: reading a credential file, running a build script, writing to a config, calling a local MCP server. All of it is invisible to a control positioned at the model boundary.

The sandbox confusion is subtler and more forgivable, because sandboxes are genuinely strong. The limit is scope rather than strength. A sandbox answers "may this command touch this path," correctly and at the kernel level. It does not answer "is this the fourth step in a sequence that started three days ago," because it has no concept of a sequence.

Runtime control infrastructure like Gensee Crate is built for that second question, and it is worth being clear that it is an addition to a sandbox rather than a replacement for one.

Why Native Agent Controls Stop Where They Do

Every major coding agent now ships credible native controls, and the fastest way to understand the runtime layer is to see precisely where those stop.

Claude Code offers six permission modes, an OS-enforced Bash sandbox using Seatbelt and bubblewrap, ordered deny rules that hold in every mode, and managed settings developers cannot override. Cursor offers three Run Modes, a Landlock and seccomp sandbox on Linux, and per-call MCP approval. Codex combines sandbox modes with an approval policy, and its cloud environments remove secrets before the agent phase begins.

These are real controls, and none of them is trying to do what a runtime layer does. Two vendors say so directly. Cursor's documentation carries the heading "Auto-review is not a security boundary," noting the classifier can allow a call you would have blocked. OpenAI documents that Codex safety monitoring runs asynchronously, that "a pause can arrive after the activity that triggered it," and that monitoring "doesn't replace sandboxing, permissions, or review of the result."

Three structural limits follow from how all of them are built:

  • They evaluate one action at a time. The unit of decision is a tool call. A sequence of individually approvable calls is approved.
  • They evaluate against a policy written beforehand. Anything the policy author did not anticipate is decided by a default or a classifier.
  • They end when the session ends. No native control documents a mechanism that carries a judgment into the next session.

Gensee's four conditions behind boundary crossings analysis is the empirical version of this argument: the conditions that compound into a crossing are visible in the sequence, not in any single step of it.

The Four Things a Runtime Layer Must Do

Strip the marketing and a runtime security layer has four jobs. A product that does two of them well is doing something useful and is not covering this category.

1. Instrument effects, not intentions. Processes started, files read and written, network connections opened, credentials accessed, and the causal links between them. The test is whether the record survives the agent disagreeing with it.

2. Produce evidence independent of the agent. A log the agent writes describes the agent's view of itself. Gensee's trace analysis of eight autonomous agent runs is a concrete demonstration: what exposed two boundary crossings was service semantics, independent evidence and persistence, not the volume of recorded activity. Volume was available in both the crossing and non-crossing runs.

3. Carry lineage across sessions. The link between a file written in one session and read in another, between a domain approved on Monday and used on Friday. Gensee's multi-session agent safety research documents the failure modes that live in this gap: context drift, state inconsistency, session-boundary confusion.

4. Enforce, at a boundary you can name. Observation without enforcement is monitoring. Enforcement requires a point where an effect can be held, and the honest question for any vendor is where that point is and what falls outside it.

Enforcement Modes and the Cost of Each

Most deployments move through three postures, and the failure characteristics differ enough that the choice matters more than it looks.

Observe. Record everything, block nothing. The only mode with no false-positive cost, and therefore the only sensible starting point. Its risk is organizational rather than technical: teams stay here indefinitely because nothing forces the next step, and an evidence trail nobody acts on is a compliance artifact rather than a control.

Warn. Surface a judgment to a human without stopping the action. Useful for calibration, and prone to a specific failure: warnings that arrive after the effect are notifications. Codex's own documentation about asynchronous monitoring is candid on this point. If the warning cannot precede the effect, it belongs to the observe posture regardless of what the interface calls it.

Block. Hold the effect. The only mode that prevents anything, and the only one with a false-positive cost paid by developers. What makes blocking survivable is scope: blocking a narrow, well-evidenced class of effect is sustainable, and blocking on a broad heuristic is not, because the team turns it off after the third bad week.

The sequencing that works is to observe until you have enough evidence to define a narrow class precisely, then block that class, then widen. The order matters because the evidence for what to block only exists after observing, and a team that blocks first is guessing.

The Performance Objection

"Runtime enforcement will slow down our agents" is the most common objection, and it deserves a straight answer rather than a reassurance.

The honest answer has three parts. First, the cost is real and it is a function of where interception happens: observing syscalls has a different profile from mediating network calls, which has a different profile from holding an effect for a decision. Second, the cost is dominated by the enforcement mode, not the instrumentation. Observation is cheap; a synchronous block that must reach a decision before an effect proceeds is where latency shows up. Third, and most importantly for a buying decision, the relevant comparison is not against an unmonitored agent. It is against the approval prompts you are already paying for, which cost a human context switch rather than milliseconds.

What should be treated with suspicion is any specific number offered without a published methodology. Overhead depends on workload shape, syscall volume, network chattiness and enforcement scope, and a figure that does not state those is not a measurement. Ask for the methodology, or treat the claim as marketing, including when it comes from a vendor you like.

Coverage Limits Worth Stating Out Loud

A category guide that lists only capabilities is not useful for evaluation. Four limits apply to runtime security as a class, whoever implements it:

  • It observes effects, not intent. It can tell you a credential file was read and where the data went. It cannot tell you whether the developer meant it.
  • Interception has a boundary, and things happen outside it. Any implementation intercepts at a specific layer. Effects produced through a channel it does not mediate are not covered, and the location of that line is the most important technical question about any product in this category.
  • Cross-session lineage requires persistence, which is itself a target. A record that spans sessions must be stored somewhere, and that store is now part of your attack surface.
  • It is downstream of a compromised host. If the machine is owned, the layer observing it is owned too.

These are not reasons to skip the layer. They are the reasons it belongs alongside native controls rather than in place of them, and any vendor unwilling to state its own version of this list is worth pressing.

Questions to Ask a Vendor

Ten questions that separate implementations, in rough order of how much they reveal:

  1. Where exactly do you intercept, and what does that miss? The answer should name a layer and a set of gaps.
  2. Is your evidence produced by the agent or beside it? Only the second survives the agent being wrong.
  3. What links an action in one session to an action in another? If nothing does, the product is session-scoped.
  4. What can you block, and at what point in the effect's lifecycle? Before, during, or after.
  5. What happens when your layer cannot start? Fail open or fail closed, and is that configurable.
  6. How does this interact with the sandbox already in place? Complement, duplicate or conflict.
  7. What is the overhead, and by what methodology was it measured? Refuse a number without a method.
  8. What does the record look like to an auditor who does not trust the agent or the vendor?
  9. How do policies get distributed and can a developer weaken them? Ask about precedence specifically.
  10. What are your coverage limits? A vendor without a ready answer has not thought about it, or does not want to.

For the tool-agnostic version of this exercise across a whole deployment rather than one product, Gensee's 15 questions to answer before deploying AI coding agents is the broader checklist.

Frequently Asked Questions

What is AI agent runtime security?

It is the layer that observes and governs an agent's actual effects as they occur, independently of the agent, and across sessions rather than within one. It sits after the permission decision and is distinguished from adjacent layers by spanning the session boundary.

How is runtime security different from an AI gateway?

A gateway sits between the agent and the model and sees API traffic. Most of what a coding agent does never crosses that path: local file reads, shell commands, local MCP calls. A runtime layer observes effects on the machine regardless of whether anything was sent to a model.

Do I still need a sandbox if I have runtime security?

Yes. A sandbox is a kernel-level boundary on what a command may touch, and nothing at a higher layer replaces that. Runtime security answers a different question, about sequences and effects across time, and the two are complementary rather than alternatives.

Does runtime enforcement slow agents down?

Instrumentation is cheap; synchronous blocking is where latency appears, because an effect waits for a decision. The comparison that matters for a rollout is against the approval prompts you are already paying for in human time. Treat any specific overhead figure without a published methodology as unverified.

Where should a team start?

Observe first, across a real workload, until you have evidence for which narrow class of effect is worth blocking. Blocking before observing means guessing at the policy, and a policy that generates false positives gets switched off.

Conclusion

The category is worth defining precisely because the alternative is buying three products that all do the same layer. Runtime security is the one that operates on effects, produces evidence independent of the agent, and crosses the session boundary. Everything else in the stack is valuable and is doing something different.

Gensee Crate is built at that layer, and Gensee's published research covers the empirical work behind it: trace analysis of controlled agent runs, reproduction of a boundary escape with open event data, and the multi-session failure modes that native controls do not span. To discuss a deployment, book a demo or see pricing.