Technology explainer
How Multi-Turn Attacks Manipulate AI Agents
Multi-turn attacks conceal one prohibited objective across a sequence of requests that may appear acceptable in isolation. Defending tool-using agents requires tracking cumulative intent, state, permissions, and action consequences across the entire workflow without blocking legitimate multi-step work.
A multi-turn attack manipulates an AI agent by spreading one prohibited objective across a sequence of requests that may look acceptable when judged separately. The agent stores context, intermediate files, tool results, and earlier decisions. An attacker exploits that continuity so each small action prepares the next, while no single message states the complete harmful plan.
The attack pattern in 30 seconds
- Hide the final intent. The conversation begins with a plausible role, harmless research question, or routine setup task.
- Build useful state. Each turn collects information, creates an artefact, changes a configuration, or establishes a premise.
- Adapt after resistance. If the agent refuses, the next request is reframed, narrowed, or routed through another tool.
- Exploit tool permissions. The risk grows when the agent can browse, execute code, modify files, send messages, or operate external services.
- Complete the objective cumulatively. Individually permitted actions combine into an outcome the system should have blocked.
Why does checking each message separately fail?
A one-turn filter asks, “Is this request harmful?” A stateful defence must also ask, “What has this conversation already done, and what outcome would the next action advance?” Those are different questions.
For example, retrieving a public fact, transforming data, and saving a file can each be legitimate. Their combination may still serve a prohibited objective, depending on the target, the data, and what the next tool call would do. The dangerous property lies in the workflow, not necessarily in one sentence.
This problem is especially important for an AI agent. A chatbot’s error may remain text on a screen. An agent can accumulate state and change an external system, so a failure at turn ten can use permissions and artefacts created during turns one through nine.
How does the manipulation develop?
- Establish a benign frame. The user presents a credible job, educational context, debugging task, or administrative workflow.
- Request a harmless component. The agent produces information or performs an action that has many legitimate uses.
- Preserve the output. The result enters memory, a file, a database, a browser session, or another tool.
- Add another component. The next request depends on earlier state but still avoids stating the complete purpose.
- Probe the boundary. A refusal reveals which wording, tool, or action triggered a safeguard.
- Adapt the route. The request changes while preserving the underlying goal, sometimes switching languages, formats, or tools.
- Trigger the consequential step. The final action makes earlier fragments operational.
This does not require the model to “forget” its rules. The agent may classify each local step incorrectly because its safety logic fails to represent the accumulated objective.
Multi-turn manipulation is not the same as prompt injection
| Risk | Where the instruction comes from | Core failure |
|---|---|---|
| Multi-turn manipulation | A user or interacting agent across a conversation | Individually plausible actions accumulate into a prohibited plan |
| Prompt injection | Often untrusted content read by the agent, such as a webpage, document, or message | External text is mistaken for an authorised instruction |
| Permission abuse | A legitimate-looking request using excessive access | The system allows an action beyond what the user or task requires |
| Memory poisoning | Incorrect or malicious state written for later use | Future decisions trust persistent information that was never properly validated |
These risks can combine. A conversation may guide the agent to open a document containing injected instructions, which then causes a privileged tool call and stores a misleading result in memory. Defence must follow data, instructions, permissions, and intent together.
Why tools increase the stakes
Tools turn language into consequences. Browsing can expose the agent to untrusted instructions. Code execution can transform or act on data. File access can reveal or modify persistent information. Messaging and payment tools can affect other people or assets.
The relevant question is not whether a tool is powerful in general. It is whether the current task needs the requested capability, on the requested target, at that moment. Least privilege limits the damage: read access does not imply write access, drafting does not imply sending, and permission for one file or recipient does not imply access to all of them.
An AI agent sandbox can restrict files, networks, processes, and credentials, but isolation alone does not decide whether an allowed action is appropriate. A safe system combines containment with policy checks and user authorisation.
What should a stateful safety system remember?
- Declared objective: what the user said they are trying to accomplish.
- Inferred workflow: the broader outcome suggested by the sequence, with uncertainty rather than a permanent accusation.
- Sensitive entities: people, accounts, systems, credentials, locations, and protected data involved.
- Completed actions: files written, services called, information retrieved, and permissions used.
- Refusals and boundary probes: repeated attempts to reach the same blocked result through different wording.
- Remaining consequences: what a proposed tool call would make possible when combined with existing state.
Memory also creates privacy and false-positive risks. Systems should retain the minimum safety-relevant state, protect it, set expiry rules, and allow legitimate users to correct a mistaken interpretation.
How can developers defend the whole workflow?
| Control | What it contributes | What it cannot solve alone |
|---|---|---|
| Conversation-level intent tracking | Connects actions across turns and reformulations | Intent inference can be uncertain or biased |
| Action-level policy checks | Evaluates the exact tool, target, arguments, and consequence | A permitted action may become harmful only in combination |
| Least-privilege credentials | Limits accessible systems and data | Does not prevent misuse inside the permitted scope |
| User confirmation | Requires explicit approval for high-impact changes | Users may approve confusing or misleading summaries |
| Rate and sequence limits | Restricts rapid accumulation or repeated probing | Slow attacks and legitimate long workflows remain difficult |
| Audit trail and rollback | Supports investigation and recovery | Some external actions cannot be reversed |
A safety guardrail should therefore operate at several layers: model response, plan, tool selection, permission boundary, and post-action monitoring.
Why confirmation prompts need context
“Allow this action?” is weak if the user cannot see what has accumulated. A useful confirmation identifies the target, data involved, external effect, whether the action is reversible, and how it relates to previous steps. It should appear close to the consequential action, not at the beginning as blanket consent for an unknown workflow.
The agent should also preserve the distinction between preparing and executing. It can draft a message without sending it, produce a proposed command without running it, or show a transaction without submitting it. These friction points give the human a meaningful chance to inspect the final composition.
How should multi-turn safety be tested?
Evaluations need adaptive conversations, not fixed prompt lists alone. A red-team system can vary persona, language, order, tool availability, and response to refusal. Tests should include benign workflows that resemble harmful ones, because blocking everything produces an unusable but superficially safe agent.
Useful measurements include prohibited-task completion, refusal timing, partial harmful progress, benign-task completion, unnecessary interruption, permission escalation, and whether the system recovers after misleading context. Human review remains important because automated judges may misread tool state or classify an incomplete sequence as success.
What the STING research adds
NewTqnia reported an EPFL framework called STING that divided harmful benchmark goals into adaptive conversational phases. Across 176 prompt instances derived from 44 scenarios, the multi-turn approach produced more illicit task completion than corresponding direct requests: AI Agents Were Easier to Manipulate One Harmless Step at a Time.
The result demonstrates a blind spot in single-prompt evaluation. It does not establish the frequency of real attacks or prove that every production agent fails similarly. The scenarios used controlled tool environments and automated judges, while deployed systems may add permissions, monitoring, and application-specific controls.
A benchmark result is a warning, not a breach statistic
Relative increases depend on the baseline, scenario, model, language, tool setup, and definition of completion. A benchmark can reveal a reproducible failure mode without predicting how often attackers encounter it in practice. Independent replication, human validation, production-like permissions, and transparent error analysis are necessary before turning one score into a claim about general agent safety.
The mental model to remember
A multi-turn attack hides the dangerous unit of analysis. The message looks harmless, while the workflow becomes harmful. Defending an agent therefore requires following the evolving objective, cumulative state, tool consequences, and permission boundaries across the whole interaction, while still allowing legitimate multi-step work.
First appeared in
AI Agents Were Easier to Manipulate One Harmless Step at a Time