Paper Page - Are Vision-Language Models Safe In The Wild? A Meme-Based Benchmark Study

VLMs are more vulnerable to harmful meme-based prompts than to synthetic images, and while multi-turn interactions offer some protection, significant vulnerabilities remain.

Rapid deployment of vision-language models (VLMs) magnifies safety risks, yet
most evaluations rely on artificial images. This study asks: How safe are
current VLMs when confronted with meme images that ordinary users share? To
investigate this question, we introduce MemeSafetyBench, a 50,430-instance
benchmark pairing real meme images with both harmful and benign instructions.
Using a comprehensive safety taxonomy and LLM-based instruction generation, we
assess multiple VLMs across single and multi-turn interactions. We investigate
how real-world memes influence harmful outputs, the mitigating effects of
conversational context, and the relationship between model scale and safety
metrics. Our findings demonstrate that VLMs show greater vulnerability to
meme-based harmful prompts than to synthetic or typographic images. Memes
significantly increase harmful responses and decrease refusals compared to
text-only inputs. Though multi-turn interactions provide partial mitigation,
elevated vulnerability persists. These results highlight the need for
ecologically valid evaluations and stronger safety mechanisms.

Source link

What's Hot

3D and 4D World Modeling: A Survey – Takara TLDR

National Gallery and Tate Have ‘Bad Blood’—and More Art News

How We Built A Unicorn Without Chasing Hype Cycles

Paper page – Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study

3D and 4D World Modeling: A Survey – Takara TLDR

EnvX: Agentize Everything with Agentic AI – Takara TLDR

P3-SAM: Native 3D Part Segmentation – Takara TLDR

National Gallery and Tate Have ‘Bad Blood’—and More Art News

Christie’s Will Auction The First Calculating Machine In History

The Art Market Isn’t Dying. The Way We Write About It Might Be.

Banksy Mural of Judge Beating Protestor Removed by Courts Service

3D and 4D World Modeling: A Survey – Takara TLDR

National Gallery and Tate Have ‘Bad Blood’—and More Art News

How We Built A Unicorn Without Chasing Hype Cycles

What's Hot

Paper page – Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study

Related Posts

Subscribe to Updates