← Back to all posts

Education & Learning

How to Build a Threat Model for AI Coding Agents

A step-by-step method covering assets, trust boundaries, entry points, and capability composition

Gensee Crate Team · · 15 min read

Security engineer reviewing runtime activity of an AI coding agent across multiple monitors

TL;DR: A threat model for AI coding agents maps every asset the agent can touch (source, secrets, CI/CD, endpoints), the trust boundaries around it, the entry points that can steer it (prompt injection is only one), and how its capabilities combine into attack paths, then pairs each path with a control. Built-in agent permission modes and sandboxes handle a narrower question, single-session behavior on one machine, so most organizations still need runtime mediation and session-spanning audit to close the gap.

An AI coding agent running with auto-accept enabled can install packages, edit files, execute shell commands, query the network, and push a branch, often without a human confirming any single step along the way. That is the default operating posture inside many engineering organizations running Claude Code, Codex, or Cursor sessions today. It is also why security teams have moved past asking whether these agents are useful and started asking how to build a threat model for AI coding agents before one of them touches a production repository. This article walks through a practical, step-by-step method: enumerate assets, map trust boundaries, catalog entry points beyond prompt injection, model how capabilities compose into real attack paths, and map each threat to a specific control.

Table of Contents

What Is a Threat Model for AI Coding Agents?

A threat model for AI coding agents is a structured inventory of what the agent can reach, who or what can influence it, and how those two facts combine into concrete attack paths, paired with a control assigned to each path. It differs from a generic security review because it treats the agent's tool access, memory, and multi-step planning as first-class attack surface, not an afterthought to the model's output quality.

It is commonly confused with two adjacent practices:

  • AI safety alignment addresses what a model says or refuses to say. A coding-agent threat model addresses what an agent with shell access, network egress, and repository write permission can actually do, regardless of how well-aligned its underlying model is.
  • A single prompt-injection test checks one input against one response. A full threat model covers an entire session, and the sessions that follow it, since a coding agent's memory, rules files, and planted artifacts can steer behavior long after the original malicious input is gone.
  • An MCP server audit reviews one entry point. Repository content, dependency metadata, fetched web pages, and agent-to-agent handoffs are equally viable entry points and need the same scrutiny.
  • Traditional STRIDE-style application threat modeling assumes a fixed, deterministic code path. A coding agent is a probabilistic planner that chooses its own sequence of actions, so the model has to account for capability composition, not just data flow.

Step 1: Inventory the Assets an Agent Can Touch

Start by listing everything the agent can read, write, or execute, because you cannot bound a risk you haven't inventoried. Most teams underestimate this list on the first pass.

At minimum, the inventory should include:

  • Source code and version control state: local working tree, branches, commit and push access, and any repositories the agent can clone or reference.
  • Secrets and credentials: API keys, cloud tokens, and service-account credentials available to the agent's process, including ones stored in configuration directories rather than environment variables.
  • CI/CD assets: workflow definitions, organization secrets, and deployment access reachable once the agent can open a pull request or merge to a protected branch.
  • The developer endpoint itself: the agent typically runs with the same OS-level permissions as the human who invoked it, so the asset list includes the whole local file system and installed toolchain, not just the project folder.
  • Agent memory and context: conversation history, cached plans, and persisted rules files that carry instructions from one session into the next.

Tip: Treat agent configuration directories (for example ~/.claude/, ~/.gemini/, ~/.codex/) as an asset class in their own right. Security researchers at Sysdig have found these directories hold API tokens, session data, and settings that are readable by any process running under the same user.

Step 2: Map the Trust Boundaries

A trust boundary marks the point where you stop assuming good faith and start requiring verification. For a coding agent, the meaningful boundaries sit between the developer, the agent process, the model provider, repository content, MCP servers, and the CI/CD pipeline.

The OWASP Secure Coding with AI Cheat Sheet draws this boundary explicitly around workflows, organization secrets, and deployment access inside CI/CD, separate from the developer-agent boundary on a workstation. That separation matters in practice: a control that only governs what happens on a laptop does nothing for an agent invoked from a GitHub Action reading a pull request.

Three boundary pairs deserve their own line in the model:

  • Developer ↔ agent: what the agent can do without a human confirming each step, and what its default permission mode actually authorizes.
  • Agent ↔ repository content: whether issue text, PR descriptions, README files, and code comments are treated as data or as instructions once the agent reads them.
  • Agent ↔ MCP servers: whether a tool's description and its returned output are validated before the agent acts on them, or trusted at face value after one-time approval.

Runtime control infrastructure like Gensee Crate sits on the agent side of these boundaries, recording which process touched which file, tool, or network destination independently of the agent's own narration.

Step 3: Catalog Entry Points Beyond Prompt Injection

Direct prompt injection through a chat message is the entry point most teams model first, and the one they stop at. It is far from the only one.

OWASP's guidance on secure coding with AI notes that, as of 2026, "issue bodies, PR descriptions, PR comments, README files, dependency changelogs, error traces, fetched web pages, and MCP tool responses all become instructions when the agent reads them." A complete entry-point catalog for a coding agent should include:

  • Repository content: comments, docstrings, and commit messages an agent parses while orienting itself in a codebase.
  • Rules and persona files: .cursorrules, CLAUDE.md, AGENTS.md, and .github/copilot-instructions.md can silently steer every future generation once modified, making them a durable, cross-session entry point rather than a one-time prompt.
  • MCP tool descriptions and responses: according to OWASP's MCP Top 10, a compromised MCP server can poison tool descriptions, shadow legitimate tool names, or alter tool definitions after a developer has already approved them, a pattern OWASP labels a rug-pull.
  • Dependencies and package metadata: changelogs and install scripts an agent reads while resolving or upgrading packages.
  • Web-fetched content: pages an agent retrieves mid-session to answer a question or check documentation.
  • Agent-to-agent handoffs: output from one agent or subagent consumed as trusted input by another, in multi-agent or subagent pipelines.

Note: NIST's evaluation work on agent hijacking found that current architectures generally combine trusted developer instructions and task-relevant data into a single unified input, which is why indirect injection through any of the entry points above can be as effective as a direct prompt. In its testing, task-specific red-team attacks raised success rates from 11% for the strongest baseline attack to 81% for the strongest new attack.

Abstract visualization of a developer workstation with multiple glowing data streams converging from repository files, terminal windows, and network connections

Step 4: Model How Capabilities Compose Into Attack Paths

An entry point only becomes dangerous once it combines with a capability the agent already holds, so this step is where isolated risks turn into attack paths. Modeling composition, not just single capabilities, is what separates a coding-agent threat model from a generic application one.

Consider a concrete chain: a dependency changelog contains a hidden instruction (entry point), the agent has shell execution and network egress (capabilities), and the invoking developer has push access to a protected branch (privilege). Individually, none of those three facts is alarming. Composed, they describe a path from a poisoned changelog to a pushed commit that exfiltrates a credential, with no single step looking anomalous in isolation.

This composition risk compounds across a session and across sessions:

  • Within a session: a planning step early in a long agent run can set up state, an installed helper script, a modified config file, a cached credential, that a later step in the same session then uses, so a reviewer who only checks the final diff misses the setup.
  • Across sessions: persistence planted in a rules file, a memory entry, or a scheduled task can lie dormant until a future, unrelated session picks it up and acts on it. Modeling this cross-session lineage, tracing a later unsafe action back to the earlier step that planted it, is what we mean by long-horizon defense, and it is where single-prompt testing structurally cannot help.

Sysdig's runtime research illustrates how fast this composition happens in practice: in one captured session, researchers observed five agent-loop iterations in ten seconds, 64 execve events, and multiple outbound HTTPS connections from a single agent process. At that speed, a human reviewing individual actions after the fact is reviewing an audit log, not making a decision in time to matter.

Why Sandboxes and Permission Modes Aren't a Threat Model

Built-in agent sandboxes and permission modes answer a narrower question than a full threat model does: they constrain what one agent process can do on one machine during one session, not what an attack path spanning entry points, sessions, and infrastructure can accomplish. That scoping is a design choice, not a flaw, but it means a threat model can't stop at the vendor's default controls.

Anthropic's own documentation for Claude Code's security model is a useful, honest illustration. As of September 2026, Manual mode starts with read-only permissions and asks before editing files or running commands, and its sandboxed Bash tool isolates filesystem and network access. Anthropic is also explicit that it "does not security-audit or manage any MCP server" a developer chooses to connect, and that trust verification for new codebases and MCP servers is disabled entirely in non-interactive mode, the mode most CI and automation pipelines actually use.

Question Built-in sandbox / permission mode Full threat model
Scope One agent process, one session, one machine All entry points, all sessions, all infrastructure the agent touches
MCP servers Vendor reviews listing criteria, not server behavior Tool descriptions and responses treated as untrusted, monitored for drift
Persistence Not tracked across sessions Cross-session lineage from planted artifact to later action
CI/automation Often runs with reduced or disabled trust checks Same mediation applied regardless of interactive vs. non-interactive mode
Evidence Session transcript Structured record of what each action actually did, independent of the agent's own narration

That last row is where mandatory mediation and effect evidence come in: a policy layer that sits outside the agent's own decision loop, checks each sensitive action against scope before it executes, and records what the action actually did rather than what the agent reported doing. Runtime control infrastructure such as Gensee Crate is built around that gap specifically, enforcing policy across a full multi-step session rather than reasoning about one prompt at a time.

Step 5: Map Threats to Controls

Every entry point and capability composition identified in Steps 1 through 4 should resolve to a named control, not a general intention to "monitor the agent." This is where the model becomes actionable.

OWASP's AI Agent Security guidance frames the underlying principle well: grant agents the minimum tools required for a specific task, scope each tool's read/write access, and for destructive, financial, administrative, or externally visible actions, separate decision-making from execution so a policy service independently validates scope and approval state before anything runs. A workable control mapping looks like this:

  • Memory and rules-file poisoning → context integrity checks. Diff and review changes to CLAUDE.md, AGENTS.md, and equivalent persona files with the same scrutiny as code, since they silently steer every future session.
  • MCP tool poisoning and rug-pulls → mandatory mediation at the tool boundary. Re-validate tool descriptions and responses at call time rather than trusting a one-time approval, consistent with OWASP's MCP guidance on shadowed tools and definitions that change after initial approval.
  • Capability composition (shell + network + credential) → task-scoped execution with rollback. Run agent work in live workspace forks where a session's changes can be inspected, then merged, promoted, rolled back, or discarded, so a composed attack path can be undone rather than merely detected after the fact.
  • Long-horizon, cross-session persistence → risk lineage. Link a later unsafe action back to the earlier session that planted it, rather than evaluating each session in isolation.
  • CI/CD exposure → sanitized context ingestion. Filter and sanitize PR titles, bodies, comments, and diffs before an automation agent includes them in its working context, since OWASP treats PR content as attacker-controlled input by default.

In practice we find the highest-leverage single change is moving policy enforcement outside the agent's own reasoning loop entirely, so a compromised or manipulated plan cannot simply talk its way past its own guardrails. That is the design intent behind pairing a sidecar enforcement layer with live workspace forks: the agent stays unmodified, no SDK rebuild required, while fork, merge, and rollback give a security team a safe way to accept the parts of a session that were legitimate and discard the parts that weren't. Our analysis of session traces across memory poisoning, prompt injection, and long-horizon attack patterns suggests that most damaging paths share the same signature: a capability composition that looked unremarkable action by action but produced an unauthorized effect once traced end to end. If your team is scoping this kind of runtime layer for existing developer environments, it's worth comparing notes with others working through the same rollout; you're welcome to book a demo or join the Discord community where these deployment patterns get discussed directly.

Abstract image of a translucent layered gate intercepting a stream of light between two nodes, suggesting a policy checkpoint

Step 6: Keep the Model Alive

A coding-agent threat model goes stale faster than most application threat models because agent capabilities, default permission modes, and MCP integrations change on a vendor release cycle you don't control. Treat it as a living document with an owner and a review trigger, not a one-time deliverable.

Set concrete review triggers:

  • A new MCP server or tool is added to any agent's approved connection list.
  • A default permission mode changes at the vendor level (verify current defaults against vendor documentation whenever you onboard a new agent version).
  • A new class of asset becomes reachable, for example when an agent gains deployment access it didn't previously have.
  • A logged action doesn't match its declared intent, which is itself a signal the model's control mapping needs updating.

Structured logging is what makes this review possible at all. OWASP recommends logging all agent decisions, tool calls, and outcomes, with structured metadata for high-risk actions covering action classification, authorization outcome, approval identifier, execution result, and policy version. Wiring that log stream into existing SIEM tooling, alongside identity and endpoint systems already in place, is usually less disruptive than it sounds, since it augments rather than replaces the controls a security team already runs.

FAQ

What are the security risks of AI coding agents?

The main risks extend well past prompt injection to include tool abuse and privilege escalation, memory poisoning through persisted rules or context files, data exfiltration via overly broad credentials, and cascading failures when one compromised action triggers a chain of further agent actions. OWASP's AI Agent Security guidance also flags excessive autonomy and approval bypass as distinct risk categories worth modeling separately from injection.

How do you secure AI coding agents?

Start by threat modeling the specific agent, environment, and integrations in use, since generic advice rarely maps to how a given team's agents actually run. From there, apply least-privilege tool scoping, sanitize untrusted context such as PR content before it reaches the agent, and add mediation outside the agent's own reasoning loop for any destructive or high-impact action.

Can AI coding agents access secrets?

Yes, often more broadly than teams expect. Agent configuration directories, environment variables, and cached credentials are frequently reachable by the agent's process because it typically runs with the same OS-level permissions as the developer who invoked it, not a restricted subset.

What is prompt injection in the context of coding agents?

Prompt injection occurs when instructions embedded in content the agent reads, an issue comment, a dependency changelog, a fetched web page, get treated as commands rather than data. NIST's research on agent hijacking found this indirect route can be highly effective because current agent architectures generally combine trusted instructions and untrusted data in a single input stream.

Are AI coding agents safe to use?

They can be used safely, but safety depends on the surrounding controls, not the agent's default configuration alone. Vendor sandboxes and permission modes reduce risk on a single machine and session; organizations running agents at scale generally need additional session-spanning audit and mediation layered on top.

Conclusion

A threat model for AI coding agents only holds up if it covers what generic security advice tends to skip: the full asset inventory, every trust boundary the agent crosses, entry points beyond a single malicious prompt, and the specific ways capabilities compose into attack paths across a session and the sessions that follow it. Built-in sandboxes and permission modes are a real, useful layer, but they answer a narrower question than the one enterprise security teams are actually asking. Closing that gap means pairing your existing identity, endpoint, and SIEM tooling with mediation and rollback that work across the full lifecycle of an agent session, not just its individual steps. If you're building this out for your own developer environments, book a demo to walk through how a sidecar deployment fits your current stack, or explore the open-source project on GitHub to see the enforcement model directly.