
TL;DR: Codex Security is OpenAI's AppSec agent, in research preview as of March 2026, that scans GitHub repositories, builds a project-specific threat model, validates findings in a sandbox, and proposes patches for human review. It reviews what's written in code before it ships; it has no visibility into what a coding agent does at runtime, such as which files it touches, which network calls it makes, or how an instruction planted in one session shapes an action in a later one. Enterprises need both layers: AppSec scanning for the code, and runtime control for the agent writing it.
Codex security means something more specific than "is my AI coding agent safe." Since March 2026, OpenAI has offered a product literally named Codex Security: an agent, still in research preview as of September 2026, that scans repositories, models what a system trusts, and flags exploitable vulnerabilities before they reach production. It's a genuine addition to the AppSec toolchain and worth understanding on its own terms.
But the name invites a mix-up with a different question entirely: what happens once the Codex coding agent (or Claude Code, or Cursor) is actually running in a developer's environment, writing files, calling tools, and carrying context across a multi-day session. This article covers what Codex Security scans and validates, where that coverage stops, and what has to sit alongside it to govern an agent's runtime behavior.
Table of Contents
- What Is Codex Security?
- How Codex Security Works: Identification, Validation, Remediation
- What the Beta Data Shows So Far
- Why Code Review Answers a Different Question Than Runtime Control
- Where Codex Security Stops: Runtime Risks Outside Its Scope
- The Complementary Layer: Runtime Control for Coding Agents
- Building a Layered AppSec and Runtime Defense Program
- FAQ
What Is Codex Security?
Codex Security is OpenAI's LLM-driven application-security agent: a research-preview product that connects to a GitHub repository, builds a project-specific threat model, and returns ranked vulnerability findings with proposed patches, according to OpenAI. It leans on language-model reasoning, test-time compute, and sandboxed validation rather than fuzzing or signature matching to decide which findings are real.
It's easy to confuse Codex Security with two adjacent things it is not:
- Not the Codex coding agent. Codex Security is a separate product built to find bugs in code. The Codex agent itself (CLI, IDE extension, or cloud form) is the tool that writes and edits code, with its own sandboxing and approval settings described in OpenAI's guidance on running Codex safely.
- Not a replacement for SAST or SCA. OpenAI's own Codex Security FAQ describes it as complementary to static analysis, adding semantic reasoning and exploit validation on top of deterministic scanners rather than replacing them.
- Not a runtime control layer. It inspects source code before or at commit time. It has no visibility into what an agent does once it starts executing tasks against live infrastructure.
That last distinction runs through the rest of this article. Runtime control infrastructure like Gensee Crate sits on the other side of it, recording and mediating what a running coding agent does to files, processes and network destinations rather than reviewing the code it wrote.
How Codex Security Works: Identification, Validation, Remediation
Codex Security runs as three sequential stages, identification, validation, and remediation, according to OpenAI's Help Center. Each stage narrows a broad codebase down to a small set of confirmed, human-reviewable findings.
Identification and threat modeling
After a scan is configured against a GitHub repository, Codex Security analyzes the codebase to understand its security-relevant structure and generates a project-specific threat model covering what the system does, what it trusts, and where it is most exposed. This step is what separates it from a rules-based scanner: it is reasoning about the shape of a specific system rather than matching known-bad patterns.
Sandboxed validation
Where possible, Codex Security pressure-tests findings in isolated, ephemeral validation environments before surfacing them, temporarily cloning the target repository to reproduce the issue and capture execution details. This is the mechanism that separates signal from noise: a candidate finding that can't be reproduced under sandbox conditions gets deprioritized rather than dumped into a backlog. For CI pipelines, teams can also drive scans through OpenAI's Codex Security CLI and TypeScript SDK rather than only through Codex web. For larger repositories, initial scans can take multiple days; later scans are typically faster because they focus on new commits and incremental changes.
Human-reviewed remediation
A confirmed finding comes with a description, file location, criticality, root cause, and a suggested patch, but the patch does not modify code automatically. It is surfaced for human review and can be turned into a pull request through a team's normal workflow. OpenAI is explicit that the product accelerates review and ranks findings, but does not replace code-level validation or human threat assessment.

What the Beta Data Shows So Far
In its first month of research preview, OpenAI reported concrete scan volume and finding counts from its beta cohort, giving an early read on how often the product finds something worth acting on.
Over that period, Codex Security scanned more than 1.2 million commits across external repositories in the beta cohort, identifying 792 critical findings and 10,561 high-severity findings, per OpenAI. Critical issues appeared in under 0.1% of scanned commits, and fourteen CVEs were assigned from the effort, with dual reporting on two of them.
Note: These figures are OpenAI's own beta-cohort telemetry as of March 2026, not an independently audited benchmark. They describe what the tool found in code that had already been written, not what happened while any agent was executing.
Why Code Review Answers a Different Question Than Runtime Control
Codex Security and runtime governance are not competing tools; they answer different questions about different moments in an agent's lifecycle, and closing one gap does not close the other.
Code review, however sophisticated, evaluates a static artifact: a commit, a pull request, a snapshot of a repository. Even continuous rescanning is still checking instructions written into files, not actions taken by a live process. It cannot see a coding agent mid-session executing shell commands, reading local files, calling an MCP tool, or hitting an internal API, because none of that activity is a line of code sitting in a repository waiting to be scanned. That gap is structural, tied to what each layer is built to look at, not something closed by a more capable scanner.
| Dimension | Codex Security (AppSec scanning) | Runtime control (agent execution) |
|---|---|---|
| What it inspects | Source code and repository structure | Live agent actions: processes, files, network, credentials |
| When it acts | Before deployment, on commits and pull requests | During the session, in real time |
| Unit of analysis | A commit or a codebase | A multi-step, multi-session agent run |
| Failure mode caught | Broken access control, injection flaws, business-logic bugs in shipped code | Memory poisoning, prompt injection, planted persistence, unauthorized tool calls |
| Output | Ranked findings with proposed patches for human review | Approve or deny decisions, fork or rollback of the agent's workspace, audit lineage |
This is the point where runtime control infrastructure like Gensee Crate picks up: not reviewing what was already written, but mediating what an agent does with it, session after session.
Where Codex Security Stops: Runtime Risks Outside Its Scope
Codex Security has no visibility into what a coding agent does once it starts executing: which processes it starts, which files or credentials it touches, which network destinations it reaches, or how instructions planted in one session shape actions in a later one.
Memory poisoning and prompt injection
An attacker can plant a malicious instruction in content the agent will later read, a README, a code comment, a dependency manifest, or a notes file the agent persists across sessions. Code scanning of the resulting output may find nothing wrong syntactically, because the vulnerability isn't in the shipped code, it's in the agent's decision path. Prompt injection carried through untrusted content, a fetched web page, a ticket description, or data returned by an MCP tool, can steer an agent toward exfiltrating secrets or editing unrelated files, none of which registers as a code vulnerability.
Long-horizon, multi-step attacks
A single prompt or commit can look entirely benign in isolation while the risk compounds across a session, or across sessions run days apart. In practice we find that the sessions most likely to cause damage are rarely the ones with an obviously malicious first prompt; they're the ones where an early, innocuous-looking action, writing a config file, registering a scheduled task, adding a helper script, only becomes dangerous several steps or several sessions later. Per-commit review evaluates each change in isolation and has no mechanism for reconstructing that lineage.
MCP tool use, credentials, and network egress
Coding agents increasingly connect to MCP servers, ticketing systems, and cloud CLIs to get work done. A mis-scoped credential or an overly permissive MCP tool call is a runtime access-control decision, not a code pattern with a fix. For the Codex coding agent (not Codex Security), OpenAI's guidance on running Codex safely points to this layer: sandbox boundaries, approval policy, and managed network egress. That is a separate discipline from scanning a repository for bugs.

The Complementary Layer: Runtime Control for Coding Agents
Runtime control governs what a coding agent is allowed to do while it is running, not what was written in the code it produced. It's the layer that has to exist alongside AppSec scanning, not instead of it.
Mandatory mediation and sidecar enforcement
Instead of rebuilding an agent's stack or waiting for a vendor SDK hook for every possible action, a sidecar sits alongside the agent and mediates its actions, file writes, shell commands, network calls, tool invocations, as they happen. This works the same way whether the agent in question is Claude Code, Codex, or Cursor, because the mediation attaches at the point where the agent touches the environment rather than inside the agent's own runtime.
Live workspace forks: fork, inspect, merge, or roll back
A coding session can be treated like a transaction rather than a stream of individual edits. That means moving an agent's work into a live workspace fork, inspecting the effect evidence it produces (the actual files created, commands run, and network calls made) before deciding to merge it, and rolling the whole workspace back cleanly if something looks wrong, instead of trying to hand-undo scattered changes after the fact.
Cross-session risk lineage
Long-horizon defense means tracing a persistence artifact planted in one session, a modified config, a new scheduled task, a quietly added dependency, to the unsafe action it enables in a later one. Our analysis suggests that treating each session as a set of reversible transactions, rather than auditing a stream of tool calls after the fact, is what makes rollback practical at the speed agents actually operate.
Products like Gensee Crate Enterprise apply this model as a sidecar attached to existing coding-agent deployments, and teams that want to inspect the mediation approach directly can start from the open-source project on GitHub. Teams already running Codex Security for code review and looking to close the runtime gap can book a demo to see how mediation attaches to a current Codex, Claude Code, or Cursor deployment.

Building a Layered AppSec and Runtime Defense Program
A mature program treats AppSec scanning and runtime control as two required layers wired into the same identity, endpoint, and monitoring stack a security team already runs, not as alternatives to choose between.
Integrating with identity, endpoint, MCP, and SIEM tooling
Enterprise coding-agent governance needs customized integration with existing identity providers for who can run scans or approve agent actions, endpoint tooling, MCP server allowlists, and centralized log ingestion. OpenAI's guidance on operating the Codex coding agent, separate from Codex Security, describes exporting agent telemetry, including prompts, tool approval decisions, tool execution results, MCP server usage, and network proxy allow or deny events, so it can be centralized in SIEM and compliance logging systems. Runtime control infrastructure needs to plug into that same telemetry path rather than creating a separate, unconnected log stream.
Where each layer reports
Runtime mediation events, approval decisions, fork or rollback actions, and cross-session lineage should land in the same identity and SIEM systems that already ingest AppSec scan findings, so a Codex Security finding on a commit and a runtime policy violation in a later agent session can be correlated by the same team rather than triaged in two disconnected tools. Details on plan-level integration options are on Gensee's pricing page.
FAQ
What is Codex Security?
Codex Security is OpenAI's research-preview AppSec agent that scans GitHub repositories, builds a project-specific threat model, and returns ranked, validated vulnerability findings with proposed patches for human review. It rolled out in research preview starting March 2026 to ChatGPT Pro, Enterprise, Business, and Edu customers, with free usage during the initial period.
How does Codex Security work?
It runs three sequential stages: identification, which builds a threat model of the repository; validation, which reproduces and pressure-tests findings in an isolated sandbox to confirm exploitability; and remediation, which surfaces a proposed patch for human review that can become a pull request. It relies on LLM reasoning and test-time compute rather than fuzzing or signature-based matching.
Does Codex Security replace SAST or manual security review?
No. OpenAI describes Codex Security as complementary to static analysis, adding semantic reasoning and exploit validation on top of existing SAST and SCA tools rather than replacing them. It accelerates review and ranks findings but does not replace code-level validation, exploitability checks, or human threat assessment.
Does Codex Security auto-apply patches?
No. Patches are surfaced for human review and can be converted into a pull request through a team's normal workflow. The tool does not modify code automatically.
Does Codex Security govern what a coding agent does at runtime?
No. Codex Security inspects source code before or at commit time and has no visibility into an agent's live actions, such as file writes, network calls, credential use, or behavior that spans multiple sessions. Governing those actions requires a separate runtime control layer.
Conclusion
Codex Security is a serious addition to the AppSec toolchain: LLM-driven scanning, project-specific threat modeling, sandboxed validation, and human-reviewed patches, with meaningful beta telemetry already logged as of its March 2026 research preview. But it answers "is this code exploitable," not "what is this agent doing right now, across this session and the ones before it." Enterprises running Codex, Claude Code, Cursor, or similar agents inside real developer environments still need a runtime layer that mediates agent actions as a sidecar, keeps agent work in live workspace forks that can be inspected, then merged, promoted, rolled back or discarded, and links planted persistence across sessions to the unsafe action it eventually enables.
If your team is scanning code but not yet mediating what the agent does with it, that gap is worth closing before it gets found for you. Book a demo or join our Discord to talk through where a sidecar control layer fits your existing Codex or Claude Code deployment.