Skip to main content

ASSURESOFT INSIGHTS

The Nearshore Advantage

Cybersecurity visual depicting AI agent architecture vulnerability review

A hacker's guide to breaking your AI agent

An attacker trying to break an AI agent would not necessarily start by jailbreaking the model. There are easier places to look, especially once the agent can read external information, call tools, access internal systems, or act with organizational credentials.

The first questions would be practical: what can the agent read, what systems does it trust, what credentials does it carry, which tools can it call, and what happens if one of its decisions can be influenced? NIST uses the term agent hijacking for attacks in which malicious instructions embedded in content an agent consumes steer its behavior.

The real trouble starts when the agent can do something with what it reads. A poisoned instruction in a runbook becomes much more serious if the same agent can edit a deployment file and send that change into CI.

Once an attacker can influence a decision, the real question is what the system allows that decision to affect. The rest of the article looks at the agent the way an attacker would, through four ways in which influence can be turned into something much more consequential. The goal is to help a reviewer find those weak points first. 

First, an attacker would find out what the agent can reach

Before attempting a compromise, an attacker would want to know how much room there is to work with. A useful starting point is the least-trusted information the agent can consume: a support ticket, a web page, a PDF, a repository comment, a database record, or a message generated by another agent.

The next question is what the agent can change after being influenced. If it can edit a repository and kick off CI, a poisoned instruction has somewhere to go.

This way of looking at the system closely resembles source-sink analysis: identify the untrusted source that can influence the agent, then look at the consequential capability the influence may eventually reach.

That exercise becomes more useful as agents gain more autonomy. OWASP describes the related problem of excessive agency as giving a model more functionality, permissions, or autonomy than the task requires.

Method 1: Poison what the agent believes

An attacker can hide an instruction within something the agent is already expected to read, such as a runbook, a support ticket, or a retrieved document. If that content reaches the model’s context, it can shape the next action without ever appearing as a direct command from the user. OWASP classifies this as indirect prompt injection.

An internal runbook includes one extra step that sends the agent to an attacker-controlled server. Everything around that line looks legitimate, so when the document is pulled into context, the planted instruction can pass as part of the procedure the agent was already following.

MITRE places this kind of manipulation within the broader class of AI agent context poisoning, in which adversarial information is inserted into the context an agent uses to reason or act.

The forged content does not have to look like a document at all. Agent frameworks usually return tool results to the model in a recognizable format, often marked by special role or tool tokens, and an attacker who can place text in the prompt may imitate it. A line claiming that a date lookup returned 2035 can arrive looking like the output of a tool the agent never called. A recent write-up on tool response injection reports this against several models, although with a deliberately small example and no published success rates, so it works better as an illustration of the mechanism than as evidence of scale.

The defensive work starts with preserving the origin of information. If the system treats a message as a tool result because it looks like one, the format has become the credential, and externally supplied content has quietly acquired the same authority as a system instruction.

Method 2: Borrow the agent’s authority

Another attack path starts with the privileges the agent already has. A user asks the agent to read customer record 123. The agent sends the request using a service account that can read every customer record in the system. If the backend checks only the service account, the user’s original restriction has effectively disappeared. 

An attacker does not need to break the backend to use this. Asking for record 124, or planting an instruction that makes the agent ask for it, is enough: the agent complies, the service account allows it, and every check along the way passes.

That is how the agent can become a confused deputy.

Research on authenticated delegation for AI agents focuses on this distinction: possessing credentials does not, by itself, establish the authority a user has delegated for a particular task.

The problem becomes easier to miss after a handoff. An agent may pass the request to another service that sees only the agent’s credentials, not the credentials of the user who originally requested record 123. By the time the action runs, the system may know what the agent is capable of doing while having lost the restriction that came with the user’s request.

Record 123 was all the user was allowed to see. That limit needs to survive every handoff. If the final service only sees the agent’s more powerful credentials, it has no way to know that accessing another record was never part of the request. 

That is also why least privilege for AI agents has to reach beyond simply giving the agent fewer permissions. Identity and task scope need to remain meaningful when the agent invokes a tool or passes work to another service.

Method 3: Combine harmless tools until they become dangerous

Some of the most interesting failures appear only when two ordinary capabilities are used together.

The agent changes the image tag in a Kubernetes deployment manifest, commits the file, and triggers CI. If the pipeline automatically deploys committed changes, the new image may reach production without any separate deployment permission being granted to the agent. Nobody granted a permission literally called deploy_to_production, yet the combination produces almost the same result.

This is where reviewing tools one by one can miss the important part. AI agent tool invocation can expose systems and capabilities that an attacker could not reach directly, and the resulting risk depends heavily on what those tools can accomplish when used together.

The same issue can appear around data. Suppose an agent can read a confidential customer file and then fetch a URL included in the task. An attacker may only need to make the agent encode part of that file in an outbound request to an attacker-controlled server. MITRE documents exfiltration through AI agent tool invocation as a distinct attack technique.

In that setup, the dangerous permission is not written anywhere as “exfiltrate data” or “deploy to production.” It appears only when the complete sequence is considered.

That changes what a useful security review looks like. The reviewer has to follow the action far enough to see what the combination actually permits, especially around deployment, external communication, money movement, access changes, or deletion.

Method 4: Make the attack survive and propagate

An attacker does not have to make the agent fail immediately. Persistence may be more useful.
A poisoned support ticket may disappear from view after the original conversation ends. If the agent turns part of it into a memory entry, however, another session may rely on that memory days later. The attacker’s instruction continues to affect the system even though the original ticket is gone. The attacker-controlled wording has crossed into a more trusted context without carrying its original warning label.

One way to describe that failure is trust laundering. The phrase is useful here as a framing device, not as established industry terminology: information that entered through an untrusted source returns later with the appearance of internal authority.

MITRE describes memory poisoning as a form of agent context poisoning in which adversarial information enters persistent memory and affects later behavior.

That changes the security problem. Long-term memory is no longer just a convenience feature. Anything written there may become input to a later decision, possibly after the original source has faded from view.

Provenance helps, but it should not be treated as a complete defense. A recent preprint on agent memory poisoning examines the limitations of relying solely on content screening and provenance ranking. 

If the agent saves a note from an external ticket and retrieves it days later, the note must retain a trace of its origin. Otherwise, an attacker may only need to poison the first interaction and let the memory system keep doing the rest. 

Persistence is only one direction. In multi-agent setups, one agent's output is another agent's input, so a poisoned note can travel through the same handoffs described in Method 2. The receiving agent sees a message from a colleague, not from a support ticket, and the attacker's wording arrives with even less of its original label. Each hop makes the content look more trusted and adds another set of permissions the attacker can borrow.

Breaking the attack before it breaks the agent

The risk becomes real when the agent can act on the poisoned input. If a manipulated document leads to a code change, a record update, or a stored note that another session later trusts, the attack has already moved beyond the prompt itself.

Together, these attack paths show that the model is only one part of the security problem. An agent can misread context, and the architecture around it determines whether that mistake ends in a bad answer, an unauthorized edit, a leaked record, or a production change.

Each method has a matching control. Against poisoned context, keep external content labeled and outside the channels the model treats as instructions: separate roles at the API level, neutralize special formatting tokens in untrusted data, and accept a tool result only if it matches a call the system actually issued. Against borrowed authority, carry the user's identity and task scope through every handoff. Against harmless tools combining, review the complete sequence rather than each permission. Against persistence and propagation, keep provenance attached to anything written to memory or passed to another agent.

A successful prompt injection gives an attacker influence. Identity, permissions, tools, persistence, and downstream trust determine what that influence is worth. For anyone reviewing an agent before it gains access to real systems, one question cuts through a surprising amount of complexity:
If this agent makes the wrong decision, what can that decision actually do?

Tags

Engineering Team

Engineering Team

AI-powered experts

A multidisciplinary team of AssureSoft engineers specializing in AI, data intelligence, product development, platform architecture, cloud, DevOps, enterprise and core systems, software quality, security, and UI/UX.

Drawing on their experience across complex technology projects, the team shares practical and technical insights on combining AI-driven productivity with human standards to help engineering leaders build software that delivers lasting value.