← Back to all posts

Education & Learning

Runtime Security vs Sandboxing for AI Coding Agents

What containers, VMs, and harness sandbox modes constrain, and what only runtime mediation can catch

Gensee Crate Team · · 14 min read

Layered glass boundaries representing runtime security governing actions across a sandboxed environment

TL;DR: Sandboxing (containers, VMs, OS-level isolation, or a coding agent's built-in harness mode) constrains where code can execute and what resources it can physically touch. Runtime security observes and governs what an agent does once it's running, mediating tool calls, tracking effect evidence, and preserving lineage across an entire multi-step session. Enterprise teams adopting AI coding agents need both: sandboxing contains a single execution boundary, while runtime security is the only layer that can authorize actions, detect long-horizon abuse, and produce an audit trail.

AI coding agents like Claude Code, Codex, and Cursor now run for extended sessions: reading repositories, calling MCP tools, editing files, and executing shell commands with minimal human review at each step. Security teams reach for sandboxing first because it is familiar: wrap the agent in a container or VM, and the blast radius of any single bad action shrinks. That instinct is correct but incomplete. Sandboxing and runtime security answer different questions, and confusing them leaves a gap that shows up specifically in long-horizon, cross-session attacks: memory poisoning, prompt injection chains, and planted persistence that only becomes dangerous several steps later.

This article compares what each actually does, where sandboxing structurally cannot follow, and how the two combine into a defense-in-depth architecture for agentic coding environments.

Table of Contents

What Is Runtime Security vs Sandboxing?

Sandboxing isolates where code runs; runtime security governs what an agent is allowed to do once it is running, mediating actions across a session and preserving evidence of every decision. They sit on different axes: one is a containment boundary, the other is continuous authorization and audit.

  • Sandboxing constrains a single execution boundary (a container, VM, or an agent harness's built-in sandbox mode) so that code cannot physically reach resources outside it.
  • Runtime security observes and mediates agent actions, tool calls, and data movement across an entire session, and links related events across sessions.
  • Sandboxing answers "can this process physically reach that file, socket, or device." Runtime security answers "should this specific action, in this specific context, be allowed right now, and can we prove what happened."
  • Sandboxes are typically scoped to one execution boundary at a time. Runtime security maintains lineage that connects an action taken in step one to a consequence that surfaces in step forty.

What Sandboxes Actually Constrain

Sandboxes constrain the physical and logical boundary a process can touch, whether that boundary is a container, a virtual machine, or a coding agent's built-in restricted mode. Each isolation model trades off strength of separation against overhead and flexibility.

Container and OS-level isolation

According to the NIST Application Container Security Guide, OS virtualization "provides a separate virtualized view of the OS to each application, thereby keeping each application isolated from all others on the server," and with containers specifically, "multiple apps share the same OS kernel instance but are segregated from each other." That shared-kernel model is efficient, but it also means a kernel-level exploit can potentially cross container boundaries in a way a hypervisor boundary resists.

VM-level isolation

NIST describes VMs as using "a hypervisor that provides hardware-level isolation of resources across VMs," where "each VM sees its own virtual hardware and includes a complete guest OS in addition to the app and its data." This is a stronger isolation model than containers alone, at the cost of higher resource overhead and slower spin-up, which matters when an agent is spawning many short-lived tool executions per session.

Harness sandbox modes in coding agents

Modern coding agents ship their own sandbox modes as a first line of containment. Anthropic's Claude Code security documentation describes a "sandboxed bash tool" that isolates commands at the filesystem and network level, and notes that in manual mode Claude Code can write only to the folder where it was started, not to parent directories, without explicit permission. OpenAI's Codex sandbox documentation frames the sandbox as "the boundary that lets the agent act autonomously without giving it unrestricted access to your machine," with a workspace-write mode that permits reading, editing within the workspace, and routine local commands, versus a danger-full-access mode that "removes the filesystem and network boundaries" entirely. OpenAI is explicit that sandboxing and approvals are separate controls: "the sandbox defines technical boundaries, the approval policy decides when the agent must stop and ask before crossing them" (as of September 2026).

Runtime control infrastructure like Gensee Crate sits alongside these harness sandbox modes and records which tool call touched which file or network destination inside the boundary they set.

Isolated developer workspace shown as a contained glass enclosure with terminal windows inside, representing container and sandbox boundaries

Where Sandboxing Structurally Stops

Sandboxing is scoped to a containment question, not an authorization or detection question, so it cannot decide whether a specific tool call should happen or reconstruct what an agent did across a session. This is not a flaw in sandboxing; it is a different job.

Palo Alto Networks states the boundary plainly: "sandboxing is not a prevention mechanism in itself. It's a containment and analysis control," and adds that it "doesn't eliminate the need for patching, segmentation, or behavior-based detection." NIST reinforces this from the operations side, noting that "traditional security solutions, such as intrusion prevention systems and web application firewalls, often do not provide suitable protection for containers" and recommends "container-aware runtime defense tools" specifically for detecting and responding to threats during runtime, separate from the isolation boundary itself. A recent agent sandbox escape shows that gap in a real incident.

The mismatch has several dimensions, which a short comparison does faster than prose:

Dimension Sandboxing (container / VM / harness mode) Runtime security
What it governs Physical/logical reach of a single execution boundary Whether a specific action is authorized, in context
Scope of visibility Inside one container, VM, or harness session Across tool calls, sessions, and identities
Tool-call authorization Coarse (allow/deny by directory, network, mode) Fine-grained (per-call policy and approval)
Cross-session memory Not tracked; each boundary is largely stateless Lineage links actions across sessions
Evidence produced Limited to what the boundary logs Effect evidence: what changed, why it was allowed
Response to abuse Terminate or rebuild the boundary Roll back, fork, or selectively undo specific actions

Sandboxing answers "can this reach that." It cannot answer "should this specific tool call, from this agent, in this session, given what happened three steps ago, be allowed." That second question is scoped to runtime security, and it is exactly the question long-horizon agent attacks are designed to exploit.

What Runtime Security Observes and Governs

Runtime security is the mediation layer that sits between an agent's decisions and the effects those decisions produce, enforcing policy on every tool call and preserving evidence of what happened. Microsoft's guidance on runtime isolation and sandboxing captures the principle directly: "each tool call should pass through a policy-enforcing mediation layer, and higher-risk actions should require stronger identity checks, approval, and auditable execution." Microsoft also recommends placing "schemas, argument validation, authorization, quotas, and approval gates in front of plugins and tools instead of exposing general shell, database, or network access," which is a governance requirement no sandbox mode alone satisfies.

In practice, this is what mandatory mediation looks like for a coding agent: every file write, shell command, and MCP tool call is evaluated against policy before it executes, not just contained after the fact. The result of each action is captured as effect evidence, meaning the system records not just that a command ran, but what it changed, what it read, and what triggered it. This is the category-first shift that runtime control infrastructure like Gensee Crate is built around: rather than replacing the sandbox layer, it adds mandatory mediation and a transactional record on top of whatever isolation boundary the agent already runs in, without requiring teams to rebuild their agent stack or fork an SDK.

That transactional layer matters because coding agent sessions are not single actions, they are chains of dozens or hundreds of steps. A transactional runtime lets a team fork an agent's workspace before a risky operation, inspect what changed, merge the safe parts, and roll back the rest, treating a multi-step coding session more like a reviewable branch than an irreversible sequence of shell commands.

Tip: When evaluating a coding agent's security posture, ask two separate questions: "what is the isolation boundary" and "what is mediating and recording tool calls inside that boundary." Vendors often answer only the first.

Long-Horizon and Cross-Session Attacks: Why Isolation Alone Misses Them

Long-horizon attacks plant a small, low-signal change early in a session and cash it in several steps or sessions later, which is exactly the pattern isolation boundaries are not designed to track. A container or VM resets or is torn down between tasks; it has no memory of what happened in a prior session, and even within one session it does not distinguish a benign file write from a planted one.

Consider a realistic sequence for a coding agent working across a multi-day feature branch. In session one, a prompt injection embedded in a fetched dependency's README instructs the agent to add an innocuous-looking helper script to a build directory. The sandbox permits this: it is a normal file write inside the workspace boundary, and nothing about it trips a container or VM control. In session two, days later, the agent (working on an unrelated task) invokes that helper script as part of a build step, and the script exfiltrates environment credentials over an allowed network path. OWASP's AI Agent Security Cheat Sheet names both halves of this pattern: "tool abuse and privilege escalation" for the first session's overly permissive write, and "data exfiltration" for the second session's use of an allowed tool call to move sensitive information out. Neither container isolation nor a harness sandbox mode connects those two events, because each session is evaluated on its own, not against a lineage of prior actions.

This is the specific gap that cross-session risk lineage closes: linking a planted persistence event to the later action that depends on it, even when the two happen in different sessions, different tool calls, or different directories. Memory poisoning follows the same shape, where a poisoned note in an agent's persistent memory or context store quietly biases a decision several steps after the poisoning occurred. OWASP's guidance to "require explicit approval for high-impact or irreversible actions" and to "log all agent decisions, tool calls, and outcomes" describes the right instinct, but a per-session sandbox has no mechanism for evaluating an action against a decision made in a different session; that correlation is a runtime security function, not an isolation function.

Abstract visualization of a multi-step chain of connected nodes representing a long-horizon agent session with a highlighted anomalous link

In our analysis of agent session traces, the actions that matter most for detection are rarely the ones that look dangerous in isolation. A single file write or a single dependency fetch looks routine on its own; it is the connection between an early, unremarkable action and a later consequential one that reveals the attack. That is a lineage problem, not a containment problem.

Why Enterprise Teams Need Both

Enterprise coding agent deployments need sandboxing and runtime security together because each closes a gap the other structurally cannot: sandboxing contains the blast radius of a single execution boundary, while runtime security governs and audits the decisions made inside it. Treating either as sufficient on its own leaves a known category of risk uncovered.

https://pub-7b00fbd2070049a3aa74950ae9141ae4.r2.dev/images/ak_4d71df17bde075ad/job_e5dd7455f20e5cfe3fb508d0a61562e4/inline-2.png

Layer one: isolation boundary

Containers, VMs, or a coding agent's own harness sandbox mode (such as Codex's workspace-write or Claude Code's sandboxed bash tool) contain what a process can physically reach. This layer should stay in place regardless of what sits above it.

Layer two: mandatory mediation

Every tool call, file write, and command inside that boundary passes through policy enforcement before it executes, with stronger identity checks and approval gates on higher-risk actions, consistent with Microsoft's mediation-layer guidance.

Layer three: live workspace fork

Agent work happens in a forkable, inspectable workspace so a team can fork before a risky step, inspect the diff, merge what's safe, and roll back what isn't, rather than treating an agent session as an all-or-nothing sequence.

Layer four: cross-session lineage

Actions are correlated across sessions and tool calls so a planted early-session change can be traced to the later action it enables, closing the specific gap long-horizon attacks are built to exploit.

In practice we find that teams who deploy sandboxing alone catch the loud failures (an agent that tries to reach a disallowed network path or write outside its workspace) but miss the quiet ones (a memory-poisoning chain or a slow-burn privilege escalation across sessions). We've seen that the combination of isolation plus mediation plus lineage is what closes both categories, not either layer alone.

Building a Layered Defense: From Laptop to CI to Production

A layered defense for AI coding agents extends the same mediation and lineage model from a developer's laptop through CI pipelines to production, rather than treating each environment as a separate security problem. Consistency across environments is what prevents a policy that holds in CI from silently disappearing on a developer's machine.

Integrating with MCP and developer tools

Model Context Protocol (MCP) has become the common integration point for coding agents reaching external tools, databases, and services, which also makes it a common path for tool abuse if calls are not mediated individually. A runtime layer that mediates MCP tool calls the same way it mediates file writes and shell commands keeps policy consistent regardless of which tool the agent is invoking.

Working with existing identity, endpoint, and SIEM tooling

Enterprise security teams do not want a parallel stack; they want agent activity to show up in the identity, endpoint, and SIEM tooling they already run. Runtime control infrastructure that operates as a sidecar to unmodified coding agents, without requiring an SDK rebuild, is designed to plug into that existing tooling rather than replace it. Gensee Crate Enterprise is built around this integration model for organizations standardizing agent governance across teams, while Gensee Crate Personal applies the same sidecar approach for individual developers who want mediation and rollback on their own machine. Teams that want to evaluate the underlying approach directly can also review the Gensee open-source repositories on GitHub.

Governing exceptions and policy changes

Policy exceptions are inevitable (a one-off script that needs broader network access, a temporary elevated credential for a migration), and the governance question is whether those exceptions are logged, time-boxed, and reviewable, or silently persist. A live workspace fork model makes this tractable: an exception is scoped to a fork, reviewed before merge, and never becomes a permanent, unaudited widening of the agent's default permissions.

If your team is comparing execution models for a coding agent rollout, walking through this with someone who works on runtime mediation daily is often faster than testing configurations in isolation. You can book a demo to see how mediation and live workspace forks apply to your specific agent stack.

FAQ

What is the difference between a sandbox and a runtime for AI agents?

A sandbox (container, VM, or harness sandbox mode) constrains where an agent's code can execute and what resources it can physically reach. Runtime security observes and governs what the agent actually does inside that boundary, mediating tool calls and preserving evidence across the session and across sessions.

Can enterprise agent systems use more than one execution model?

Yes, and most production agent architectures do. A common pattern uses container or VM isolation as the containment boundary while a runtime security layer mediates tool calls, tracks lineage, and enforces policy inside and across those boundaries.

Is Claude Code sandboxed by default?

As of Anthropic's current security documentation, Claude Code offers a sandboxed bash tool with filesystem and network isolation, and in manual mode restricts writes to the starting folder and its subfolders without explicit permission for parent directories. Whether sandboxing is the default behavior depends on the mode and version in use, so teams should confirm current defaults against Anthropic's documentation directly.

How does MCP integrate with sandbox environments?

MCP tool calls typically execute within whatever isolation boundary the coding agent or its host environment provides, but the sandbox boundary alone does not evaluate whether an individual MCP call should be authorized. That authorization and audit function belongs to a runtime mediation layer sitting alongside the sandbox.

What isolation model is most secure for AI agent sandboxes executing untrusted code?

There is no single answer independent of the workload: VM-level isolation with a hypervisor provides stronger separation than shared-kernel containers for untrusted code, according to NIST's container security guidance, but at higher resource cost. Most enterprise deployments choose based on the sensitivity of the workload and layer runtime mediation on top regardless of which isolation model they pick.

Conclusion

Sandboxing and runtime security solve different problems, and enterprise teams running AI coding agents need both rather than treating one as a substitute for the other. Sandboxing, whether through containers, VMs, or an agent's own harness mode, contains what a process can physically reach. Runtime security mediates every tool call, preserves effect evidence, and links actions across an entire multi-step session, which is the only layer equipped to catch memory poisoning, prompt injection, and long-horizon attacks that unfold across steps a single sandbox boundary never sees together.

If your organization is standardizing how coding agents get deployed across development teams, it's worth mapping your current isolation setup against the mediation and lineage gaps described here. Explore more detail in our FAQ, or book a demo to walk through how a sidecar runtime layer fits alongside the sandboxing you already have. For teams that prefer to start in the community first, our Discord is a good place to compare notes with other engineers working on the same problem.