What Is Agentjacking and Why AI Coding Teams Should Care

In June 2026, researchers demonstrated agentjacking: fake Sentry bug reports could hijack AI coding agents and make them execute malicious commands through trusted MCP integrations.

Content

Make Your Applications Secure Today

Sign up for a personalized demo to see how DerScanner can meet your Application Security needs

In June 2026, researchers at Tenet Security published a proof of concept with an uncomfortable premise: a fake bug report can take over an AI coding agent. Not through malware, not through phishing, not through any compromise of the victim's machine. Through a Sentry error event that the agent reads as trusted diagnostic guidance and acts on (Tenet Security's research). They called the technique agentjacking, and in testing it worked against the coding agents most teams actually run.

 

The attack, step by step

The entry point is a Sentry DSN, the Data Source Name that applications use to submit error events. Sentry treats DSNs as public and write-only by design; they sit in frontend JavaScript on production websites and turn up by the thousand through GitHub search. Tenet identified at least 2,388 organizations with exposed DSNs that accepted injected events.

With a DSN, an attacker can POST an arbitrary error event into a project: message, stack trace, tags, breadcrumbs, all attacker-controlled. Tenet formatted these events to look exactly like Sentry's own remediation guidance, including an instruction to run a diagnostic npm package.

A developer asks their coding agent to fix unresolved Sentry issues, a completely routine request. The agent queries Sentry through its MCP server, receives the injected event alongside real ones, and has no way to tell them apart, because the MCP server returns everything as trusted system output. The agent follows the "remediation steps" and executes the attacker's package with the developer's full privileges. Text in, remote code execution out. In testing against more than 100 real-world targets, the technique succeeded 85% of the time across Claude Code, Cursor, and Codex, including the agents of a Fortune 100 company.

 

Why this differs from classic prompt injection

Prompt injection usually means hostile text inside content the model was asked to read: a poisoned webpage or a document. Agentjacking targets something more structural. The vector is the MCP trust boundary itself: the convention that whatever a connected tool returns is system data rather than untrusted input. The attacker never talks to the model. They talk to a legitimate telemetry platform, and the platform's own integration delivers the payload with trusted framing attached.

Sentry responded by deploying a content filter for the specific payload, and Tenet shipped a mitigation tool called Agent-JackStop, but filtering known strings at the ingestion layer is a patch on an architectural issue. Any MCP-connected data source that accepts external input, ticketing systems, log aggregators, code review comments, inherits the same problem.

Research cited by CSA Labs found that 43% of sampled MCP server implementations contained command injection flaws and 30% allowed unrestricted URL fetching. The pattern reaches first-party servers too: in April 2026 Microsoft patched CVE-2026-32211, a missing-authentication flaw in Azure MCP Server rated CVSS 9.1. OWASP now classifies MCP tool poisoning as a recognized attack category, which is standards-speak for "stop treating this as theoretical." We look at where it landed in the new agentic risk taxonomy in OWASP Agentic.

 

What mitigation actually looks like

Three controls do most of the work, and none of them require abandoning agents.

Approval gates on execution. An agent may propose commands; a human confirms anything that runs code, touches credentials, or leaves the sandbox. The 85% success rate above was measured largely against agents running in permissive modes, where confirmation is switched off for convenience.

Context isolation. Data returned by external tools should enter the model marked as untrusted content, separated from instructions. Agents that render tool output and system guidance in one undifferentiated stream are exactly the agents that followed Tenet's fake remediation steps.

Scoped tool permissions. An agent connected to Sentry for triage does not need shell access in the same session. Narrow tool sets per task shrink what a hijacked agent can do, the same least-privilege logic applied to a new kind of identity.

 

Where static analysis fits

There is a fourth control that tends to get overlooked because it is unfashionably ordinary: the MCP servers and agent integrations themselves are code. TypeScript, Python, Go. They parse untrusted input, build shell commands, fetch URLs, and hold tokens. The 43% command injection figure above describes ordinary software vulnerabilities in ordinary software, and SAST finds those the same way it finds them anywhere else: taint tracking from tool input to command execution, hardcoded credentials, unvalidated fetch destinations.

The practical move is to pull agent plumbing into the application security perimeter. Scan the MCP servers a team writes or forks. Run SCA on agent framework dependencies, which have already become a supply chain target in their own right, as the Miasma campaign showed on the npm side. Treat a new MCP integration the way a new public API endpoint gets treated, because from an attacker's perspective that is what it is.

Agentjacking is a preview of the next few years: the agent layer stitches together components that were each individually secure and produces a system that is not. The components can still be scanned. If MCP servers and agent integration code are in the codebase right now and have never been through static analysis, a DerScanner demo with that exact code is a concrete place to start, and the SAST overview explains what the analysis covers.

Loading blogs...
Get Started

Ready to Reduce Technical Debt and
Improve Security?

Clean code. Fewer risks. Stronger software

dashboard