← Back to all posts

Education & Learning

Runtime Security vs PR Review: Closing the Coverage Gap

Why code review can't see what an AI agent does before it commits code

Gensee Crate Team · · 13 min read

Translucent security layer monitoring an AI coding agent's live workflow

TL;DR: PR review, human or AI, evaluates the diff an agent proposes to merge; it cannot see the commands, file reads, or network calls the agent made while producing that diff. Runtime security mediates and records those actions as they happen, across the full multi-step session, so risks like memory poisoning or a planted persistence mechanism aren't invisible until someone stumbles on the code weeks later. The two controls are complementary, not competing: pairing session-level runtime enforcement with diff-level review, through a fork, inspect, merge, and rollback workflow, closes the coverage gap that either one leaves alone.

A pull request from an AI coding agent can look completely clean: passing tests, a tight diff, a clear commit message. It can still hide the fact that the agent read a credentials file, called an external API, or left a script behind that never shows up in a GitHub diff view. That gap, what runtime security catches versus what PR review catches, is the distinction enterprise teams adopting AI coding agents like Claude Code, Codex, and Cursor need to get right before they treat code review as their security control.

This article walks through what each approach actually inspects, why PR review is structurally scoped to a different question than runtime enforcement, and how the two combine into a layered workflow that covers both the artifact and the process that produced it.

Table of Contents

What Is Runtime Security vs. PR Review?

Runtime security and PR review answer different questions asked at different points in an AI coding agent's work. PR review, whether performed by a human or an AI reviewer bot, examines the code that lands in a commit; runtime security mediates and records what the agent actually did, tool call by tool call, while producing that commit. They are not rival controls competing for the same job; they cover non-overlapping evidence.

The distinction that trips teams up is assuming an AI-powered review tool is already "runtime" security because it's automated and fast. It isn't. A few things separate them:

  • Timing: PR review happens after the work is done, before merge. Runtime security operates continuously, during the session.
  • Evidence: PR review sees the diff. Runtime security sees commands executed, files read or written (inside and outside the repo), network calls, and MCP tool invocations.
  • Actor: PR review is a human or an AI bot commenting on a pull request. Runtime security is enforcement infrastructure that mediates the agent's actions as they happen.
  • Goal: PR review judges whether the final code is correct and safe. Runtime security judges whether the process that created it was safe.

What Pull Request Review Actually Catches

Pull request review, at its core, inspects a bounded set of file changes and decides whether they're safe to merge. It was never designed to observe the process that produced those changes.

GitHub's own documentation describes the workflow plainly: reviewers examine "commits, file changes, and diffs," look at each changed file individually, leave comments on specific lines, and submit a verdict of Comment, Approve, or Request Changes. That's the full scope of the input: the diff, and whatever surrounding code the reviewer chooses to open.

AI-assisted review has scaled this same pattern rather than changed its shape. Cloudflare has described launching "up to seven specialized reviewers" against every merge request, covering security, performance, code quality, documentation, release management, and internal compliance checks. Over one measured month (March 10 to April 9, 2026), Cloudflare's system completed 131,246 review runs across 48,095 merge requests in 5,169 repositories, with a median review time of about three minutes forty seconds. Cloudflare is explicit that this "isn't a replacement for human code review," and notes its own reviewers "see the diff and surrounding code" but lack full context for architectural intent.

Wiz draws the same boundary from a security-tooling angle: automated code review, including static analysis, secrets detection, and dependency scanning, runs against source and configuration "before runtime." Wiz contrasts this directly with IAST (interactive application security testing), which "monitors code behavior in runtime for vulnerabilities." Everything in the PR review category, human or AI, operates on committed artifacts. None of it operates on the live process that created them.

The Bridge: Why Review Can't See the Session

Code review is structurally scoped to the diff because it was built to answer "is this proposed patch safe and correct," a question that assumes a bounded change submitted by a person. An autonomous coding agent doesn't just produce a patch: across a session it can run shell commands, read files inside and outside the repository, call MCP tools, reach the network, and carry state across dozens of turns. None of that has to appear as a line in a diff.

Runtime control infrastructure like Gensee Crate records that session-level activity, which commands ran and which files and network destinations the agent touched, as evidence kept separate from the diff.

As of September 2026, in the Claude Code documentation, Anthropic itself flags this boundary problem: it recommends sandboxing for filesystem and network enforcement because Bash deny and ask rules cover only "the command text Claude writes" and are explicitly "not a security boundary around the program." Reviewing the commands an agent reports running is not the same as enforcing what it is able to do.

PR / Code Review Runtime Security
Question it answers Is this diff safe to merge? Was this session's behavior safe?
Unit of evidence The diff Commands, file reads/writes, network calls, MCP tool calls
When it acts After the work is done, pre-merge While the agent is working
Scope The repository Repository, file system, MCP tools, network
Typical catch Insecure pattern, bad logic, missing tests Secret exfiltration, out-of-repo reads, unauthorized commands, planted persistence
Blind spot Actions that never reach the diff Architectural fit, long-term code quality

Runtime Security Controls for AI Coding Agents

Runtime security closes exactly the gap the table above describes: it treats the agent's session, not the eventual pull request, as the thing to secure. The operating principle is mandatory mediation: every consequential action, a shell command, a file write, an outbound request, an MCP tool call, has to pass through an enforcement point before it executes, rather than being reconstructed afterward from logs.

This is the gap a newer category of runtime control infrastructure for coding agents is built to address. Platforms such as Gensee Crate are designed to sit alongside unmodified agents like Claude Code, Codex, and Cursor as a sidecar, mediating actions without requiring a rebuilt SDK or a modified agent runtime.

Abstract visualization of a security sidecar layer monitoring a coding agent's workflow

Three mechanics separate this from reviewing logs after the fact:

  • Sidecar enforcement: the control point runs alongside the agent process itself, mediating tool calls as they're issued rather than sampling command history later.
  • Transactional runtime: the session is treated like a transaction, something that can be forked, inspected, merged, or rolled back, instead of a one-way stream of actions.
  • Effect evidence: instead of trusting an agent's own narration of what it did, the control layer records the actual effects (files touched, network destinations, processes spawned), which is what makes long-horizon defense possible across a full session rather than a single prompt.

Note: None of this requires replacing PR review. It adds a second, independent source of evidence about the same unit of work: the session that produced the diff, not just the diff itself.

We've seen agent sessions where a single early tool call, fetching a dependency's README, introduces instructions that only manifest as a state change dozens of turns later. That's why session-level effect evidence, not a single prompt's output, has to be the unit of analysis for anything claiming to defend against long-horizon risk. The enforcement core behind this pattern is also available as an open source project on GitHub for teams that want to inspect the sidecar mechanics directly.

Long-Horizon Risk: How Attacks Unfold Across a Session

Long-horizon risk means the plant and the payoff happen at different points in a session, or in different sessions entirely, so no single prompt or diff line reveals the attack. Understanding the mechanics matters because it explains why single-prompt filters and diff review both miss it.

The OWASP Top 10 for Agentic Applications 2026, published December 9, 2025 and developed with more than 100 industry experts, researchers, and practitioners, frames the underlying problem directly: agentic systems "plan, act, and make decisions across complex workflows," which means security guidance has to be built around multi-step behavior rather than isolated requests.

Memory Poisoning

An agent's persistent memory or context store gets polluted with attacker-controlled content, often planted in a README, a ticket description, or a dependency's documentation. The agent later treats that planted content as a trusted instruction, sometimes many turns after it was first read.

Prompt Injection Across Turns

An injected instruction doesn't have to fire on the turn it's introduced. It can sit dormant in context or memory until a later turn, or a later session on the same repository, triggers it. Single-exchange injection filters only look at one prompt-response pair; they have no visibility into a payload that separates the plant from the use by dozens of turns or several days.

Cross-Session Persistence and MCP

An agent legitimately connected to an MCP server for tool integration can receive a compromised or malicious tool response that plants a credential, a cron entry, or a config change during one session. The payoff, exfiltration or lateral movement, happens in a later session, sometimes run by a different developer or a scheduled job, with no direct diff connecting the two events. Cross-session risk lineage is designed specifically to trace that later unsafe action back to the earlier tool call that planted it, even though a diff review of either session in isolation would show nothing unusual.

Developer workstation showing a coding agent terminal with a layered timeline of session activity

In practice we find that the sessions worth the closest scrutiny are rarely the ones with one obviously malicious command. They're the ones where an early, unremarkable action, a routine file write or a config edit, only becomes dangerous in light of what happens dozens of turns later.

Combining Runtime Defense and PR Review

Runtime security and PR review aren't a choice between two controls; used together they form a layered workflow that covers both the process and the product. In practice, the two layers connect through a four-step cycle: fork, inspect, merge, or roll back.

  1. Fork: the agent's workspace is forked before exploratory or higher-risk work begins, isolating the session from the main branch and from any shared state.
  2. Inspect: the runtime layer surfaces effect evidence and session lineage for the forked work, commands run, files touched, anything flagged as persistence, alongside the diff that PR review evaluates on its own terms.
  3. Merge: once both the code (via PR review) and the session (via runtime evidence) clear, the change merges into the main branch normally.
  4. Roll back: if either the diff or the session evidence raises a concern, the forked workspace, and any side effects created inside it, can be rolled back without ever touching the main branch or shared infrastructure.

https://pub-7b00fbd2070049a3aa74950ae9141ae4.r2.dev/images/ak_4d71df17bde075ad/job_cbedab7c0c6a0d03ada48e216e96fcbe/inline-2.png

Our analysis suggests the highest-value point to apply this pairing is exactly where the two forms of evidence overlap: a merge decision backed by both a clean diff and a clean session trail is a materially different risk than one backed by the diff alone. Teams evaluating this workflow for individual use can start with Gensee Crate Personal; teams rolling it out across a codebase and multiple agents typically move to Gensee Crate Enterprise. If it would help to see the fork, inspect, merge, and rollback cycle against a real repository and agent setup, you can book a demo.

Enterprise Governance: Plugging Into Existing Tooling

Enterprise adoption of coding agents rarely starts from a blank slate: security teams already run identity providers, endpoint agents, MCP gateways, and a SIEM collecting everything else. Runtime security for coding agents needs to plug into that stack, not ask for a parallel one.

In practice, that means a few specific integration points:

  • Identity: runtime events tie back to a real employee or service account, not just a hostname or a session ID.
  • Endpoint: runtime security shares signal with existing endpoint controls rather than duplicating a separate agent on the same machine.
  • MCP: tool calls through MCP servers are treated as a distinct governed surface, since an overly permissive or compromised MCP server is itself an attack path, not just a convenience integration.
  • SIEM: session-level evidence exports into the SIEM security teams already query, in a format that fits existing detection rules and dashboards.

Gensee Crate Enterprise is built around this integration model for organization-wide deployment across existing identity, endpoint, MCP, and SIEM tooling. Details on plans and how they scale by team size are on the pricing page, and engineering teams comparing implementation options are welcome to join the Discord to ask questions directly.

FAQ

Does AI-powered code review replace the need for runtime security?

No. AI code review, like human review, still evaluates the diff after the work is done; it has no visibility into commands, file reads, or network calls made during the session that produced the diff. Runtime security covers that separate layer of evidence, so the two work best paired rather than treated as substitutes.

Can PR review catch prompt injection or memory poisoning attacks?

PR review can sometimes catch the visible symptom, an odd line of code or a suspicious dependency, but not the underlying mechanism. Prompt injection and memory poisoning often plant an instruction that only produces a visible effect many turns or sessions later, well outside what a single diff review examines.

What is a live workspace fork in AI agent security?

A live workspace fork treats an agent's session like a transaction: it can be forked before risky work begins, inspected for what actually happened, then merged if the code and session both check out or rolled back if either raises a concern. This lets teams contain and reverse agent work without it ever touching shared infrastructure.

Do runtime security tools require changing my coding agent or its SDK?

Sidecar-based runtime security is designed to mediate an unmodified agent's actions from alongside the agent process, not from inside a rebuilt SDK. Agents like Claude Code, Codex, and Cursor can run as they normally would, with the sidecar enforcing and recording actions in real time.

How is memory poisoning different from prompt injection?

Prompt injection is an attacker-controlled instruction delivered through an input the agent processes, such as a file, webpage, or tool response, during a session. Memory poisoning is a related but distinct problem where that planted content persists in the agent's memory or context store, so it can influence behavior in a later turn or an entirely separate session.

Conclusion

PR review and runtime security aren't competing answers to the same question. PR review, human or AI, judges whether a diff is safe to merge; runtime security mediates and records what an agent actually did to produce it, across the full session, including the commands, file reads, and network calls that never reach the diff at all. Long-horizon risks like memory poisoning, dormant prompt injection, and cross-session persistence live precisely in that gap, which is why enterprise teams adopting coding agents need both layers, not one dressed up to look like the other.

If your team is running Claude Code, Codex, Cursor, or similar agents across developer environments and wants session-level enforcement alongside your existing review process, book a demo to see how it fits your identity, endpoint, and SIEM setup, or join our Discord to talk through the architecture with the team building it. You can also browse the blog for more on long-horizon agent risk and transactional defense patterns.