A summary of safety-related benchmarks for agents
| Method | Year | Evaluation Focus | #Tasks/Records | Target LLMs/Agents |
|---|---|---|---|---|
| Simulation-based Benchmarks | ||||
| BIPIA (Yi et al., 2023) | 2023 | IPI Attacks | 5 Scen., 250 Goals | GPT-3.5, GPT-4, etc. |
| ToolEmu (Chan et al., 2024) | 2023 | Emulated Tool Risks | 36 Tools, 144 Cases | GPT-4, LLaMA-2-70B |
| InjecAgent (Zhan et al., 2024) | 2024 | Tool-Integrated IPI | 17 User Tools, 62 Attacker Tools, 1,054 Cases | Qwen, Mistral, etc. |
| AgentDojo (Debenedetti et al., 2024) | 2024 | Third-Party Instructions | 97 Tasks, 629 Cases | Gemini-1.5-Flash, Claude-3-Sonnet, etc. |
| AgentHarm (Andriushchenko et al., 2024) | 2024 | Harmful Behaviors | 110 Tasks, 11 Cats | GPT-4o, Claude-3.5, etc. |
| RedCode (Guo et al., 2024a) | 2024 | Code Vulnerabilities | 4k+ Cases, 25 Types | GPT-4o, Claude-3.5, etc. |
| VPI-Bench (Cao et al., 2024) | 2024 | Visual Prompt Injections | 306 Cases, 5 Platforms | GPT-4o, Claude-3.5, Gemini-1.5-Pro, etc. |
| R-Judge (Yuan et al., 2024) | 2024 | Risk Identification (Logs) | 569 Recs, 27 Scen. | GPT-3.5/4o, LLaMA-3-8B, etc. |
| SALAD-Bench (Shao et al., 2024a) | 2024 | Hierarchical Safety (MCQ) | 21k Samples, 16 tasks, 66 Cats. | GPT-4, Claude-3-Sonnet, etc. |
| h4rm3l (Draguns et al., 2024) | 2024 | Jailbreak Attack Synthesis | 2 656 Attacks | GPT-4o, Claude-3.5, etc. |
| SG-Bench (Zhang et al., 2024r) | 2024 | Safety Generalization | 1,442 Queries, 6 Cats | GPT-4, Claude-3-Sonnet, etc. |
| ChemSafetyBench (Li et al., 2024t) | 2024 | Chemistry Safety | 30k Samples, 3 Tasks | GPT-4o, Claude-3.5, etc. |
| ToolSword (Ye et al., 2024a) | 2024 | Tool-Use Safety | 6 Scen., 3 Stages | GPT-4, Claude-3.5, etc. |
| PrivacyLens (Shao et al., 2024b) | 2024 | Privacy Norm Awareness | 493 Seeds/Vignettes/Trajectories | GPT-4, Claude-3-Sonnet, etc. |
| Real-Interaction Benchmarks | ||||
| SafeBench (Guo et al., 2022) | 2022 | Driving Safety | 8 Scen., 100 Routes, 2,352 Cases | 4 RL Algs, 4 Input Types |
| ASB (Zhang et al., 2024f) | 2024 | Attack–Defense (10 Scen.) | 400+ Tools | GPT-4o, Claude-3.5, etc. |
| SafeAgentBench (Yin et al., 2024) | 2024 | Embodied Hazards | 750 Tasks | GPT-4, LLaMA-3-8B, etc. |
| Agent-SafetyBench (Zhang et al., 2024s) | 2024 | Safety Risks (8 Risk Cats) | 349 Envs, 2 000 Cases | GPT-4o, Claude-3.5, etc. |
| AdvWeb (Liu et al., 2024p) | 2024 | Adversarial Robustness (Web) | 200 Target Tasks | GPT-4V, Gemini-1.5-Pro |
| ST-WebAgentBench (Shlomov et al., 2024) | 2024 | Web Safety / Trust | 222 Tasks (Each with ST Policies) | Open-Source Agents |
| Dissecting Adversarial (Liu et al., 2024p) | 2024 | Multimodal Robustness | 200 Adversarial Tasks | GPT-4V, Gemini-1.5-Pro |
| Haicosystem (Zhou et al., 2024f) | 2024 | Human-AI Sandbox (92 Scen.) | 1,840 Sims | SOTA LLMs |
| ARE (Wu et al., 2024a) | 2024 | Adversarial Robustness (Graph) | 200 Targeted Tasks | GPT-4V, Gemini-1.5-Pro, etc. |
| WASP (Evtimov et al., 2025) | 2025 | Web Safety (Adversarial) | 84 Tasks, 42 Scen. (2 Envs) | GPT-4o, Claude-3.5 |
| Refusal-Trained LLMs (Kumar et al., 2025) | 2025 | Browser Jailbreaking | 100 Harm Behaviors | GPT-4o, o1-preview |
| SafeArena (Lee et al., 2025b) | 2025 | Web-Agent Misuse | 500 Tasks (Safe/Harmful) | GPT-4o, Claude-3.5, etc. |
| OpenAgentSafety (Vijayvargiya et al., 2025) | 2025 | Real-World Safety (8 Cats) | 350+ Multi-Turn Tasks | Claude-3.5, o1-mini |
| Method | Year | Evaluation Focus | #Tasks/Records | Target LLMs/Agents |
|---|---|---|---|---|
| 2023 | 5 Scen., 250 Goals | GPT-3.5, GPT-4, etc. | ||
| ToolEmu ( | 2023 | Emulated Tool Risks | 36 Tools, 144 Cases | GPT-4, LLaMA-2-70B |
| InjecAgent ( | 2024 | Tool-Integrated | 17 User Tools, 62 Attacker Tools, 1,054 Cases | Qwen, Mistral, etc. |
| AgentDojo ( | 2024 | Third-Party Instructions | 97 Tasks, 629 Cases | Gemini-1.5-Flash, Claude-3-Sonnet, etc. |
| AgentHarm ( | 2024 | Harmful Behaviors | 110 Tasks, 11 Cats | GPT-4o, Claude-3.5, etc. |
| RedCode ( | 2024 | Code Vulnerabilities | 4k+ Cases, 25 Types | GPT-4o, Claude-3.5, etc. |
| VPI-Bench ( | 2024 | Visual Prompt Injections | 306 Cases, 5 Platforms | GPT-4o, Claude-3.5, Gemini-1.5-Pro, etc. |
| R-Judge ( | 2024 | Risk Identification (Logs) | 569 Recs, 27 Scen. | GPT-3.5/4o, LLaMA-3-8B, etc. |
| SALAD-Bench ( | 2024 | Hierarchical Safety ( | 21k Samples, 16 tasks, 66 Cats. | GPT-4, Claude-3-Sonnet, etc. |
| h4rm3l ( | 2024 | Jailbreak Attack Synthesis | 2 656 Attacks | GPT-4o, Claude-3.5, etc. |
| SG-Bench ( | 2024 | Safety Generalization | 1,442 Queries, 6 Cats | GPT-4, Claude-3-Sonnet, etc. |
| ChemSafetyBench ( | 2024 | Chemistry Safety | 30k Samples, 3 Tasks | GPT-4o, Claude-3.5, etc. |
| ToolSword ( | 2024 | Tool-Use Safety | 6 Scen., 3 Stages | GPT-4, Claude-3.5, etc. |
| PrivacyLens ( | 2024 | Privacy Norm Awareness | 493 Seeds/Vignettes/Trajectories | GPT-4, Claude-3-Sonnet, etc. |
| SafeBench ( | 2022 | Driving Safety | 8 Scen., 100 Routes, 2,352 Cases | 4 |
| 2024 | Attack–Defense (10 Scen.) | 400+ Tools | GPT-4o, Claude-3.5, etc. | |
| SafeAgentBench ( | 2024 | Embodied Hazards | 750 Tasks | GPT-4, LLaMA-3-8B, etc. |
| Agent-SafetyBench ( | 2024 | Safety Risks (8 Risk Cats) | 349 Envs, 2 000 Cases | GPT-4o, Claude-3.5, etc. |
| AdvWeb ( | 2024 | Adversarial Robustness (Web) | 200 Target Tasks | GPT-4V, Gemini-1.5-Pro |
| ST-WebAgentBench ( | 2024 | Web Safety / Trust | 222 Tasks (Each with | Open-Source Agents |
| Dissecting Adversarial ( | 2024 | Multimodal Robustness | 200 Adversarial Tasks | GPT-4V, Gemini-1.5-Pro |
| Haicosystem ( | 2024 | Human-AI Sandbox (92 Scen.) | 1,840 Sims | |
| 2024 | Adversarial Robustness (Graph) | 200 Targeted Tasks | GPT-4V, Gemini-1.5-Pro, etc. | |
| 2025 | Web Safety (Adversarial) | 84 Tasks, 42 Scen. (2 Envs) | GPT-4o, Claude-3.5 | |
| Refusal-Trained LLMs ( | 2025 | Browser Jailbreaking | 100 Harm Behaviors | GPT-4o, o1-preview |
| SafeArena ( | 2025 | Web-Agent Misuse | 500 Tasks (Safe/Harmful) | GPT-4o, Claude-3.5, etc. |
| OpenAgentSafety ( | 2025 | Real-World Safety (8 Cats) | 350+ Multi-Turn Tasks | Claude-3.5, o1-mini |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.