Table 9.

A summary of attacks and defenses for agents

Attack/DefenseMethodYearCategorySubcategoryAccess to LLMsDataset
I Indirect Prompt Injection: Malicious Instruction and Jailbreak
Greshake et al. (2023) 2023Indirect Prompt InjectionPrompt InjectionBlack-BoxCustom (webpage injections)
Wu et al. (2024c)2024Indirect Prompt InjectionPrompt InjectionBlack-BoxCustom (website prompts)
TensorTrust (Toyer et al., 2023)2023Indirect Prompt InjectionPrompt InjectionBlack-BoxSubmissions Dataset
Perez and Ribeiro (2022) 2022Indirect Prompt InjectionPrompt InjectionBlack-BoxOpenAI Examples
AttackHOUYI (Liu et al., 2023e)2023Indirect Prompt InjectionPrompt InjectionBlack-BoxCustom (multilingual prompts)
Pedro et al. (2023) 2023Indirect Prompt InjectionPrompt InjectionBlack-BoxCustom (Langchain apps)
Zhan et al. (2025) 2025Indirect Prompt InjectionJailbreakWhite-BoxCustom (tool outputs)
PANDORA (Deng et al., 2024b)2024Indirect Prompt InjectionJailbreakBlack-BoxCustom (document embeddings)
Imprompter (Fu et al., 2024a)2024Indirect Prompt InjectionJailbreakWhite-BoxCustom (LeChat, ChatGLM)
Instruction Hierarchy (Wallace et al., 2024)2024IPI DefensePrivilege ManagementWhite-BoxCustom (message privilege levels)
Instruction Detection (Wen et al., 2025a)2025IPI DefenseDetectionWhite-BoxPEEP Dataset
DefenseInstruction Detection and Removal (Chen et al., 2025d)2025IPI DefenseDetectionWhite-BoxCrafted datasets
Task Shield (Jia et al., 2024a)2024IPI DefenseDetectionWhite-BoxAgentDojo
FATH (Wang et al., 2024g)2024IPI DefenseAuthenticationBlack-BoxCustom (response tagging)
Spotlighting (Hines et al., 2024)2024IPI DefensePrompt EngineeringBlack-BoxCustom (provenance marking)
Design Pattern (Beurer-Kellner et al., 2025)2025IPI DefenseSecure DesignBlack-BoxCustom (factory floor)
II Component-Level: Short-Term Memory, Long-Term Memory, Tool Integration and MCP Server
Contextual Backdoor (Liu et al., 2024a)2024Memory AttackBackdoorBlack-BoxAgentBench, WebShop
Watch out for your agents (Yang et al., 2024h)2024Memory AttackBackdoorBlack-BoxLLM-based agents
DemonAgent (Zhu et al., 2025)2025Memory AttackBackdoorWhite-BoxAgentBench, WebArena
AttackBadAgent (Wang et al., 2024q)2024Memory AttackBackdoorWhite-BoxAgentInstruct, Mind2Web
AgentPoison (Chen et al., 2024d)2024Memory AttackBackdoorBlack-BoxHotpotQA, AgentBench
TrojanRAG (Pan et al., 2024)2023Memory AttackBackdoorWhite-BoxNQ, HotpotQA
BadRAG (Xue et al., 2024a)2024Memory AttackBackdoorWhite-BoxNQ, TriviaQA
Phantom (Chaudhari et al., 2024)2024Memory AttackBackdoorWhite-BoxNQ, TriviaQA, SQuAD
BreakingAgent (Zhang et al., 2024a)2024Memory AttackPrompt InjectionBlack-BoxCustom (Gmail agents)
AttackMINJA (Dong et al., 2025b)2025Memory AttackPrompt InjectionBlack-BoxWebshop, MIMIC-III, eICU, MMLU
PoisonedRAG (Zou et al., 2024b)2024Memory AttackKnowledge PoisoningBlack-BoxNQ, HotpotQA, MS-MARCO
Corpus Poisoning (Zhong et al., 2023)2023Memory AttackKnowledge PoisoningWhite-BoxNatural Questions, MS-MARCO
Context Extension (Luo et al., 2023; Geng et al., 2024)2023Memory DefenseContext ExtensionBlack-BoxCustom (long context benchmarks)
Prompt Leakage Defense (Agarwal et al., 2024)2024Memory DefenseDetectionBlack-BoxMulti-turn (4 domains)
DefenseAgentSafe (Mao et al., 2025)2025Memory DefenseSecure DesignWhite-BoxCustom (multi-agent tasks)
TrustRAG (Zhou et al., 2025b)2025Memory DefenseDetectionWhite-BoxCustom (RAG datasets)
Astute RAG (Wang et al., 2024a)2024Memory DefenseDetectionBlack-BoxSQuAD 2.0
RobustRAG (Xiang et al., 2024a)2024Memory DefenseSecure DesignWhite-BoxCustom (RAG datasets)
UDora (Zhang et al., 2025e)2025Tool ManipulationJailbreakWhite-BoxAgentHarm
ToolSword (Ye et al., 2024a)2024Tool ManipulationRed TeamingBlack-BoxToolBench, AgentBench
ToolCommander (Wang et al., 2024f)2024Tool ManipulationAdversarial AttackBlack-BoxToolBench
WIPI (Wu et al., 2024b)2024Tool ManipulationPrompt InjectionBlack-BoxChatGPT Web Agents, Web GPTs
AutoCMD (Jiang et al., 2025)2025Tool ManipulationAdversarial AttackBlack-BoxLangChain, KwaiAgents, QwenAgent
MPMA (Wang et al., 2025h)2025MCP ManipulationAdversarial AttackWhite-BoxCustom (MCP servers)
AttackFerrag et al. (2025) 2025MCP ManipulationAdversarial AttackBlack-BoxCIAQA, AgentBackdoorEval, etc.
 Kong et al. (2025) 2025MCP ManipulationRed TeamingGrey-BoxCustom (Claude, Filesystem, Chroma, Gmail)
AgentGuard (Chen and Cong, 2025)2025Tool & MCP DefenseRed TeamingWhite-BoxCustom (tool orchestration)
DefensePrivacyAsst (Zhang et al., 2024t)2024Tool & MCP DefenseSecure DesignBlack-BoxCustom (tool-using agents)
MCIP (Jing et al., 2025)2025Tool & MCP DefenseMaclious DetectionWhite-BoxCustom (MCIP-Bench)
GuardAgent (Xiang et al., 2024c)2024Tool & MCP DefenseSecure DesignBlack-BoxEICU-AC, Mind2Web-SC
Adversarial Multimodal Injection (Bagdasaryan et al., 2023b)2023VLM AttackPrompt InjectionWhite-BoxCustom (perturbed media)
Zhang (2024) 2024VLM AttackPrompt InjectionBlack-BoxOSWorld, VisualWebArena
Wu et al. (2024a)2024VLM AttackPrompt InjectionBlack-BoxVisualWebArena
Fu et al. (2023b)2023VLM AttackAdversarial AttackWhite-BoxCustom (VLM + tool calls)
EIA (Liao et al., 2024)2024VLM AttackPrompt InjectionBlack-BoxMind2Web
AdvAgent (Xu et al., 2024)2024VLM AttackRed TeamingBlack-BoxCustom (web tasks)
Ma et al. (2024c)2024VLM AttackAdversarial AttackBlack-BoxCustom (simulated GUI)
AttackFine-Print Injections (Chen et al., 2025a)2025VLM AttackPrompt InjectionBlack-BoxCustom (234 adversarial webpages)
SmoothVLM (Xu et al., 2024c)2024VLM DefenseDetectionWhite-BoxCustom (adversarial datasets)
BlueSuffix (Li et al., 2024g)2024VLM DefenseDetectionBlack-BoxCustom (VLM benchmarks)
LlavaGuard (Helff et al., 2024)2024VLM DefenseDetectionWhite-BoxCustom (multimodal safety)
DefenseJailDAM (Chen et al., 2024a)2024VLM DefenseDetectionBlack-BoxCustom (jailbreak detection)
III System-Level: Multi-Agent System and Embodied Agent
Prompt Infection (Lee and Tiwari, 2024)2024Multi-Agent AttackPrompt InjectionBlack-BoxCustom (interconnected ecosystems)
Morris-II (Cohen et al., 2024)2024Multi-Agent AttackPrompt InjectionBlack-BoxCustom (GenAI apps)
Ju et al. (2024) 2024Multi-Agent AttackPrompt InjectionWhite-BoxCustom (multi-agent communities)
AgentSmith (Gu et al., 2024)2024Multi-Agent AttackJailbreakMixedCustom (multimodal agents)
CORBA (Zhou et al., 2025e)2025Multi-Agent AttackCommunication AttackMixedAutoGen, Camel (various topologies)
Agent-in-the-Middle (He et al., 2025)2025Multi-Agent AttackCommunication AttackBlack-BoxAutoGen, MetaGPT, ChatDev
Evil Geniuses (Tian et al., 2023)2023Multi-Agent AttackJailbreakBlack-BoxCAMEL, MetaGPT, ChatDev
The Wolf Within (Tan et al., 2024)2024Multi-Agent AttackInfection AttackBlack-BoxMLLM societies
AttackX-Teaming (Rahman et al., 2025)2025Multi-Agent AttackJailbreakWhite-BoxHarmBench
AutoDefense (Zeng et al., 2024c)2024Multi-Agent DefenseDetectionBlack-BoxCurated harmful prompts, DAN, Stanford Alpaca
PsySafe (Zhang et al., 2024w)2024Multi-Agent DefenseFrameworkBlack-BoxCustom (Camel, AutoGen)
APOSG Framework (Standen et al., 2025)2025Multi-Agent DefenseFrameworkWhite-BoxCustom (DPA, RADE)
LLAMOS (Lin et al., 2024a)2024Multi-Agent DefenseDetectionBlack-BoxGLUE datasets
Audit-LLM (Song et al., 2024a)2024Multi-Agent DefenseDetectionBlack-BoxCERT r4.2, CERT r5.2, PicoDomain
DefenseXGuard-Train (Rahman et al., 2025)2025Multi-Agent DefenseTrainingBlack-BoxCustom (30K jailbreaks)
PVEP (Cheng et al., 2024a)2024Embodied Agent AttackAdversarial AttackMixedVIMA
Fime et al. (2025) 2025Embodied Agent AttackAdversarial AttackWhite-BoxCustom
Wang et al. (2025d)2025Embodied Agent AttackAdversarial AttackMixedBridgeData V2, LIBERO
RoboPair (Robey et al., 2024b)2024Embodied Agent AttackJailbreakMixedCustom
BadRobot (Zhang et al., 2025c)2025Embodied Agent AttackJailbreakBlack-BoxCode as Policies, ProgPrompt, VoxPoser, VisProg
POEX (Lu et al., 2025)2025Embodied Agent AttackJailbreakMixedHarmful-RLBench
Liu et al. (2024b)2024Embodied Agent AttackBackdoorBlack-BoxProgPrompt, VoxPoser, VisProg
BALD (Jiao et al., 2025)2025Embodied Agent AttackBackdoorMixedProgPrompt, VoxPoser, VisProg
EAI (Li et al., 2024l)2024Embodied Agent AttackRed TeamingBlack-BoxVirtualHome, BEHAVIOR
HASARD (Tomilin et al., 2025)2025Embodied Agent AttackRed Teaming-Custom
HEAL (Chakraborty et al., 2025)2025Embodied Agent AttackRed TeamingBlack-BoxVirtualHome, BEHAVIOR
ERT (Karnik et al., 2025)2025Embodied Agent AttackRed TeamingBlack-BoxCalbin, RLBench
AttackX-ICM (Zhou et al., 2025c)2025Embodied Agent AttackRed TeamingBlack-BoxAGNOSTOS
EAD (Wu et al., 2024f)2024Embodied Agent DefenseDetectionBlack-BoxCustom
GPSR (Shirasaka et al., 2024)2024Embodied Agent DefenseFrameworkBlack-BoxEAsafetyBench, SafeAgentBench
Pinpoint (Wang et al., 2025b)2025Embodied Agent DefenseDetectionWhite-BoxEAsafetyBench, SafeAgentBench
DefenseSafeVLA (Zhang et al., 2025a)2025Embodied Agent DefenseTrainingWhite-BoxSafety-CHORES
IV Agentic Attack & Defenses
Fang et al. (2024b)2024Agentic AttackExploitationGray-BoxCustom (CVE vulnerability)
HPTSA (Zhu et al., 2024c)2024Agentic AttackExploitationBlack-BoxCustom (zero-day vulnerability)
AutoAdvExBench (Carlini et al., 2025)2025Agentic AttackDefense ExploitationMixedCustom (defense papers)
RedAgent (Xu et al., 2024a)2024Agentic AttackJailbreakingBlack-BoxGPT Applications, HarmBench
ALI-Agent (Wang et al., 2024c)2024Agentic AttackSafety EvaluationBlack-BoxCustom (alignment scenarios)
AttackAutoRedTeamer (Zhou et al., 2025a)2025Agentic AttackRed TeamingMixedHarmBench, Custom benchmarks
Shieldagent (Chen et al., 2025e)2025Agentic DefenseFrameworkBlack-BoxCustom (multi-agent interactions)
TrustAgent (Yang et al., 2024b)2024Agentic DefenseFrameworkBlack-BoxCustom (multi-domain tasks)
DefenseAegisLLM (Cai, 2025)2025Agentic DefenseFrameworkBlack-BoxWMDP, StrongReject

or Create an Account

Close Modal
Close Modal