A summary of attacks and defenses for VLMs
| Attack/Defense | Method | Year | Category | Subcategory | Target models | Datasets |
|---|---|---|---|---|---|---|
| Adversarial Attack | Caption Attack (Schlarmann and Hein, 2023) | 2023 | White-box | Task-specific+V | OpenFlamingo | MS-COCO/Flickr30k/OK-VQA/VizWiz |
| VisBreaker (Cui et al., 2024b) | 2023 | White-box | Task-specific+V | LLaVA/BLIP-2/ InstructBLIP | MS-COCO/VQA V2/ScienceQA-Image/TextVQA/ POPE/MME | |
| CroPA (Luo et al., 2024a) | 2024 | White-box | Cross-prompt+VL | OpenFlamingo/ BLIP-2/ InstructBLIP | MS-COCO/VQA-v2 | |
| GroundBreaker (Gao et al., 2024b) | 2024 | White-box | Task-specific+V | MiniGPT-v2 | RefCOCO/RefCOCO+/ RefCOCOg | |
| Stop-reasoning Attack (Wang et al., 2024r) | 2024 | White-box | CoT attack+V | MiniGPT-4/ OpenFlamingo/ LLaVA | ScienceQA/A-OKVQA | |
| InstructTA (Wang et al., 2023b) | 2023 | Gray-box | Encoder attack+V | BLIP-2/ InstructBLIP/ MiniGPT-4/LLaVA/CogVLM | ImageNet-1K/LLaVA-Instruct-150K/MS-COCO | |
| Attack Bard (Dong et al., 2023) | 2023 | Black-box | Transfer-based+V | Bard/GPT-4V/Bing Chat/ERNIE Bot | NeurIPS’17 adversarial competition dataset | |
| AttackVLM (Zhao et al., 2024f) | 2024 | Black-box | Transfer-based+V | BLIP/UniDiffuser/ Img2Prompt/BLIP-2/LLaVA/MiniGPT-4 | ImageNet-1K/MS-COCO | |
| DynVLA (Gu et al., 2025) | 2025 | Black-box | Transfer-based+VL | InstructBLIP/MiniGPT4/ LLaVA/Gemini | MS-COCO/VQA-v2 | |
| AdvDiffVLM (Guo et al., 2024b) | 2024 | Black-box | Generator-based+V | MiniGPT-4/LLaVA/UniDiffuser/MiniGPT-4/BLIP/BLIP-2/Img2LLM | NeurIPS’17 adversarial competition dataset/MS-COCO | |
| AnyAttack (Zhang et al., 2024j) | 2024 | Black-box | Generator-based+V | CLIP/BLIP/BLIP2/ InstructBLIP/ MiniGPT-4 | MSCOCO/Flickr30K/ SNLI-VE | |
| CAVALRY-V (Zhang et al., 2025d) | 2025 | Black-box | Generator-based+V | GPT-4.1/Gemini/ QwenVL/InternVL/ LLaVA/Aria/ MiniCPM | MMBench-Video/Video-MME | |
| Latency-Energy Attack | Verbose Images (Gao et al., 2024c) | 2024 | White-box | Task-specific+V | BLIP/BLIP2/ InstructBLIP/ MiniGPT-4 | MS-COCO/ImageNet |
| Jailbreak Attack | Image Hijack (Bailey et al., 2023) | 2023 | White-box | Target-specific+V | LLaVA | Alpaca training set/AdvBench |
| Adversarial Alignment Attack (Carlini et al., 2024b) | 2024 | White-box | Target-specific+V | MiniGPT-4/LLaVA/LLaMA Adapter | toxic phrase dataset | |
| VAJM (Qi et al., 2024a) | 2024 | White-box | Universal attack+V | MiniGPT-4/LLaVA/ InstructBLIP | VAJM training set/ VAJM test set/ RealToxicityPrompts | |
| imgJP (Niu et al., 2024) | 2024 | White-box | Universal attack+V | MiniGPT-4/MiniGPT-v2/LLaVA/InstructBLIP/mPLUG-Owl2 | AdvBench-M | |
| UMK (Wang et al., 2024i) | 2024 | White-box | Universal attack+VL | MiniGPT-4 | AdvBench/VAJM training set/VAJM test set/ RealToxicityPrompts | |
| HADES (Li et al., 2024u) | 2024 | White-box | Hybrid method+V | LLaVA/GPT-4V/Gemini-Pro-Vision | HADES dataset | |
| Jailbreak in Pieces (Shayegani et al., 2023) | 2023 | Black-box | Transfer-based+V | LlaVA /LLaMA-Adapter V2 | Jailbreak in Pieces dataset | |
| Figstep (Gong et al., 2025) | 2023 | Black-box | Manual pipeline+V | LLaVA-v1.5/MiniGPT-4/CogVLM/GPT-4V | SafeBench | |
| SASP (Wu et al., 2023) | 2023 | Black-box | Prompt leakage+L | LLaVA/GPT-4V | Celebrity face image dataset/CelebA/LFWA | |
| VRP (Ma et al., 2024b) | 2024 | Black-box | Manual pipeline+V | LLaVA/Qwen-VL-Chat/ OmniLMM /InternVL Chat-V1.5/Gemini-Pro-Vision | RedTeam-2k/HarmBench | |
| HIMRD (Teng et al., 2024) | 2024 | Black-box | Manual pipeline+VL | LLaVA/DeepSeek/ GPT-4o/Gemini/ Qwen-VL | SafeBench/tiny-SafeBench | |
| IDEATOR (Wang et al., 2025c) | 2025 | Black-box | Red teaming+VL | LLaVA/InstructBLIP/MiniGPT-4 | AdvBench/VAJM test set | |
| Prompt Injection Attack | Adversarial Prompt Injection (Bagdasaryan et al., 2023a) | 2023 | White-box | Optimization-based+V | LLaVA/PandaGPT | Self-collected dataset |
| Typographic Attack (Qraitem et al., 2024) | 2024 | Black-box | Typography-based+V | LLaVA/MiniGPT4/ InstructBLIP/GPT-4V | OxfordPets / StanfordCars / Flowers / Aircraft / Food101 | |
| Backdoor & Poisoning Attack | Shadowcast (Xu et al., 2024f) | 2024 | Poisoning | Tuning-stage+VL | LLaVA/MiniGPT-v2/InstructBLIP | cc-sbu-align dataset |
| Instruction-Tuned Backdoor (Liang et al., 2024a) | 2024 | Backdoor | Tuning-stage+VL | OpenFlamingo/BLIP-2/LLaVA | MIMIC-IT/COCO/Flickr30K | |
| Anydoor (Lu et al., 2024a) | 2024 | Backdoor | Testing-stage+VL | LLaVA/MiniGPT-4/InstructBLIP/BLIP-2 | VQAv2/SVIT/DALL-E dataset | |
| BadVLMDriver (Ni et al., 2024) | 2024 | Backdoor | Tuning-stage+V | LLaVA/MiniGPT-4 | nuScenes dataset | |
| ImgTrojan (Tao et al., 2024) | 2024 | Backdoor | Tuning-stage+VL | LLaVA | LAION | |
| Jailbreak Defenses | JailGuard (Zhang et al., 2023f) | 2023 | Detection | Detection+VL | GPT-3.5/MiniGPT-4 | Self-collected dataset |
| GuardMM (Sharma et al., 2024a) | 2024 | Detection | Detection+V | GPT-4V/LLAVA/MINIGPT-4 | Self-collected dataset | |
| AdaShield (Wang et al., 2024p) | 2024 | Prevention | Prevention+V | LLaVA/CogVLM/MiniGPT-v2 | Figstep/QR | |
| MLLM-Protector (Pi et al., 2024) | 2024 | Prevention | D+P+V | Open-LLaMA/LLaMA/LLaVA | Safe-Harm-10K | |
| ECSO (Gou et al., 2024) | 2024 | Prevention | Prevention+V | LLaVA/ShareGPT4V/mPLUG-OWL2/Qwen-VL-Chat/InternLM-XComposer | MM-SafetyBench/VLSafe/ VLGuard | |
| InferAligner (Wang et al., 2024h) | 2024 | Prevention | Prevention+VL | LLaMA2/LLaVA | AdvBench/TruthfulQA/ MM-Harmful Bench | |
| BlueSuffix (Zhao et al., 2025) | 2024 | Prevention | Prevention+VL | LLaVA/MiniGPT-4/Gemini | MM-SafetyBench/RedTeam-2k | |
| DPS (Zhou et al., 2025d) | 2025 | Prevention | Prevention+V | Qwen-VL-Plus/GPT-4o/Gemini-1.5-Flash | RTA-100/MultiTrust/Self-Gen/MM-SafetyBench/HADES/ VisualAttack | |
| ETA (Ding et al., 2025a) | 2025 | Prevention | Prevention+VL | LLaVA/InternVL/InternLM-XComposer/LLaMA3.2-Vision | SPA-VL/MM-SafetyBench/FigStep |
| Attack/Defense | Method | Year | Category | Subcategory | Target models | Datasets |
|---|---|---|---|---|---|---|
| Adversarial Attack | Caption Attack ( | 2023 | White-box | Task-specific+V | OpenFlamingo | MS-COCO/Flickr30k/OK-VQA/VizWiz |
| VisBreaker ( | 2023 | White-box | Task-specific+V | LLaVA/BLIP-2/ InstructBLIP | MS-COCO/VQA V2/ScienceQA-Image/TextVQA/ POPE/MME | |
| CroPA ( | 2024 | White-box | Cross-prompt+VL | OpenFlamingo/ BLIP-2/ InstructBLIP | MS-COCO/VQA-v2 | |
| GroundBreaker ( | 2024 | White-box | Task-specific+V | MiniGPT-v2 | RefCOCO/RefCOCO+/ RefCOCOg | |
| Stop-reasoning Attack ( | 2024 | White-box | CoT attack+V | MiniGPT-4/ OpenFlamingo/ LLaVA | ScienceQA/A-OKVQA | |
| InstructTA ( | 2023 | Gray-box | Encoder attack+V | BLIP-2/ InstructBLIP/ MiniGPT-4/LLaVA/CogVLM | ImageNet-1K/LLaVA-Instruct-150K/MS-COCO | |
| Attack Bard ( | 2023 | Black-box | Transfer-based+V | Bard/GPT-4V/Bing Chat/ERNIE Bot | NeurIPS’17 adversarial competition dataset | |
| AttackVLM ( | 2024 | Black-box | Transfer-based+V | BLIP/UniDiffuser/ Img2Prompt/BLIP-2/LLaVA/MiniGPT-4 | ImageNet-1K/MS-COCO | |
| DynVLA ( | 2025 | Black-box | Transfer-based+VL | InstructBLIP/MiniGPT4/ LLaVA/Gemini | MS-COCO/VQA-v2 | |
| AdvDiffVLM ( | 2024 | Black-box | Generator-based+V | MiniGPT-4/LLaVA/UniDiffuser/MiniGPT-4/BLIP/BLIP-2/Img2LLM | NeurIPS’17 adversarial competition dataset/MS-COCO | |
| AnyAttack ( | 2024 | Black-box | Generator-based+V | CLIP/BLIP/BLIP2/ InstructBLIP/ MiniGPT-4 | MSCOCO/Flickr30K/ SNLI-VE | |
| CAVALRY-V ( | 2025 | Black-box | Generator-based+V | GPT-4.1/Gemini/ QwenVL/InternVL/ LLaVA/Aria/ MiniCPM | MMBench-Video/Video-MME | |
| Latency-Energy Attack | Verbose Images ( | 2024 | White-box | Task-specific+V | BLIP/BLIP2/ InstructBLIP/ MiniGPT-4 | MS-COCO/ImageNet |
| Jailbreak Attack | Image Hijack ( | 2023 | White-box | Target-specific+V | LLaVA | Alpaca training set/AdvBench |
| Adversarial Alignment Attack ( | 2024 | White-box | Target-specific+V | MiniGPT-4/LLaVA/LLaMA Adapter | toxic phrase dataset | |
| 2024 | White-box | Universal attack+V | MiniGPT-4/LLaVA/ InstructBLIP | |||
| imgJP ( | 2024 | White-box | Universal attack+V | MiniGPT-4/MiniGPT-v2/LLaVA/InstructBLIP/mPLUG-Owl2 | AdvBench-M | |
| 2024 | White-box | Universal attack+VL | MiniGPT-4 | AdvBench/VAJM training set/VAJM test set/ RealToxicityPrompts | ||
| 2024 | White-box | Hybrid method+V | LLaVA/GPT-4V/Gemini-Pro-Vision | |||
| Jailbreak in Pieces ( | 2023 | Black-box | Transfer-based+V | LlaVA /LLaMA-Adapter V2 | Jailbreak in Pieces dataset | |
| Figstep ( | 2023 | Black-box | Manual pipeline+V | LLaVA-v1.5/MiniGPT-4/CogVLM/GPT-4V | SafeBench | |
| 2023 | Black-box | Prompt leakage+L | LLaVA/GPT-4V | Celebrity face image dataset/CelebA/LFWA | ||
| 2024 | Black-box | Manual pipeline+V | LLaVA/Qwen-VL-Chat/ OmniLMM /InternVL Chat-V1.5/Gemini-Pro-Vision | RedTeam-2k/HarmBench | ||
| HIMRD ( | 2024 | Black-box | Manual pipeline+VL | LLaVA/DeepSeek/ GPT-4o/Gemini/ Qwen-VL | SafeBench/tiny-SafeBench | |
| IDEATOR ( | 2025 | Black-box | Red teaming+VL | LLaVA/InstructBLIP/MiniGPT-4 | AdvBench/VAJM test set | |
| Prompt Injection Attack | Adversarial Prompt Injection ( | 2023 | White-box | Optimization-based+V | LLaVA/PandaGPT | Self-collected dataset |
| Typographic Attack ( | 2024 | Black-box | Typography-based+V | LLaVA/MiniGPT4/ InstructBLIP/GPT-4V | OxfordPets / StanfordCars / Flowers / Aircraft / Food101 | |
| Backdoor & Poisoning Attack | Shadowcast ( | 2024 | Poisoning | Tuning-stage+VL | LLaVA/MiniGPT-v2/InstructBLIP | cc-sbu-align dataset |
| Instruction-Tuned Backdoor ( | 2024 | Backdoor | Tuning-stage+VL | OpenFlamingo/BLIP-2/LLaVA | MIMIC-IT/COCO/Flickr30K | |
| Anydoor ( | 2024 | Backdoor | Testing-stage+VL | LLaVA/MiniGPT-4/InstructBLIP/BLIP-2 | VQAv2/SVIT/DALL-E dataset | |
| BadVLMDriver ( | 2024 | Backdoor | Tuning-stage+V | LLaVA/MiniGPT-4 | nuScenes dataset | |
| ImgTrojan ( | 2024 | Backdoor | Tuning-stage+VL | LLaVA | ||
| Jailbreak Defenses | JailGuard ( | 2023 | Detection | Detection+VL | GPT-3.5/MiniGPT-4 | Self-collected dataset |
| GuardMM ( | 2024 | Detection | Detection+V | GPT-4V/LLAVA/MINIGPT-4 | Self-collected dataset | |
| AdaShield ( | 2024 | Prevention | Prevention+V | LLaVA/CogVLM/MiniGPT-v2 | Figstep/QR | |
| MLLM-Protector ( | 2024 | Prevention | D+P+V | Open-LLaMA/LLaMA/LLaVA | Safe-Harm-10K | |
| 2024 | Prevention | Prevention+V | LLaVA/ShareGPT4V/mPLUG-OWL2/Qwen-VL-Chat/InternLM-XComposer | MM-SafetyBench/VLSafe/ VLGuard | ||
| InferAligner ( | 2024 | Prevention | Prevention+VL | LLaMA2/LLaVA | AdvBench/TruthfulQA/ MM-Harmful Bench | |
| BlueSuffix ( | 2024 | Prevention | Prevention+VL | LLaVA/MiniGPT-4/Gemini | MM-SafetyBench/RedTeam-2k | |
| 2025 | Prevention | Prevention+V | Qwen-VL-Plus/GPT-4o/Gemini-1.5-Flash | RTA-100/MultiTrust/Self-Gen/MM-SafetyBench/HADES/ VisualAttack | ||
| 2025 | Prevention | Prevention+VL | LLaVA/InternVL/InternLM-XComposer/LLaMA3.2-Vision | SPA-VL/MM-SafetyBench/FigStep |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.