Table 10.

A summary of safety-related benchmarks for agents

MethodYearEvaluation Focus#Tasks/RecordsTarget LLMs/Agents
Simulation-based Benchmarks
BIPIA (Yi et al., 2023)2023IPI Attacks5 Scen., 250 GoalsGPT-3.5, GPT-4, etc.
ToolEmu (Chan et al., 2024)2023Emulated Tool Risks36 Tools, 144 CasesGPT-4, LLaMA-2-70B
InjecAgent (Zhan et al., 2024)2024Tool-Integrated IPI17 User Tools, 62 Attacker Tools, 1,054 CasesQwen, Mistral, etc.
AgentDojo (Debenedetti et al., 2024)2024Third-Party Instructions97 Tasks, 629 CasesGemini-1.5-Flash, Claude-3-Sonnet, etc.
AgentHarm (Andriushchenko et al., 2024)2024Harmful Behaviors110 Tasks, 11 CatsGPT-4o, Claude-3.5, etc.
RedCode (Guo et al., 2024a)2024Code Vulnerabilities4k+ Cases, 25 TypesGPT-4o, Claude-3.5, etc.
VPI-Bench (Cao et al., 2024)2024Visual Prompt Injections306 Cases, 5 PlatformsGPT-4o, Claude-3.5, Gemini-1.5-Pro, etc.
R-Judge (Yuan et al., 2024)2024Risk Identification (Logs)569 Recs, 27 Scen.GPT-3.5/4o, LLaMA-3-8B, etc.
SALAD-Bench (Shao et al., 2024a)2024Hierarchical Safety (MCQ)21k Samples, 16 tasks, 66 Cats.GPT-4, Claude-3-Sonnet, etc.
h4rm3l (Draguns et al., 2024)2024Jailbreak Attack Synthesis2 656 AttacksGPT-4o, Claude-3.5, etc.
SG-Bench (Zhang et al., 2024r)2024Safety Generalization1,442 Queries, 6 CatsGPT-4, Claude-3-Sonnet, etc.
ChemSafetyBench (Li et al., 2024t)2024Chemistry Safety30k Samples, 3 TasksGPT-4o, Claude-3.5, etc.
ToolSword (Ye et al., 2024a)2024Tool-Use Safety6 Scen., 3 StagesGPT-4, Claude-3.5, etc.
PrivacyLens (Shao et al., 2024b)2024Privacy Norm Awareness493 Seeds/Vignettes/TrajectoriesGPT-4, Claude-3-Sonnet, etc.
Real-Interaction Benchmarks
SafeBench (Guo et al., 2022)2022Driving Safety8 Scen., 100 Routes, 2,352 Cases4 RL Algs, 4 Input Types
ASB (Zhang et al., 2024f)2024Attack–Defense (10 Scen.)400+ ToolsGPT-4o, Claude-3.5, etc.
SafeAgentBench (Yin et al., 2024)2024Embodied Hazards750 TasksGPT-4, LLaMA-3-8B, etc.
Agent-SafetyBench (Zhang et al., 2024s)2024Safety Risks (8 Risk Cats)349 Envs, 2 000 CasesGPT-4o, Claude-3.5, etc.
AdvWeb (Liu et al., 2024p)2024Adversarial Robustness (Web)200 Target TasksGPT-4V, Gemini-1.5-Pro
ST-WebAgentBench (Shlomov et al., 2024)2024Web Safety / Trust222 Tasks (Each with ST Policies)Open-Source Agents
Dissecting Adversarial (Liu et al., 2024p)2024Multimodal Robustness200 Adversarial TasksGPT-4V, Gemini-1.5-Pro
Haicosystem (Zhou et al., 2024f)2024Human-AI Sandbox (92 Scen.)1,840 SimsSOTA LLMs
ARE (Wu et al., 2024a)2024Adversarial Robustness (Graph)200 Targeted TasksGPT-4V, Gemini-1.5-Pro, etc.
WASP (Evtimov et al., 2025)2025Web Safety (Adversarial)84 Tasks, 42 Scen. (2 Envs)GPT-4o, Claude-3.5
Refusal-Trained LLMs (Kumar et al., 2025)2025Browser Jailbreaking100 Harm BehaviorsGPT-4o, o1-preview
SafeArena (Lee et al., 2025b)2025Web-Agent Misuse500 Tasks (Safe/Harmful)GPT-4o, Claude-3.5, etc.
OpenAgentSafety (Vijayvargiya et al., 2025)2025Real-World Safety (8 Cats)350+ Multi-Turn TasksClaude-3.5, o1-mini

or Create an Account

Close Modal
Close Modal