The Safety Reminder: A Soft Prompt To Reactivate Delayed Safety Awareness In Vision-Language Models

arXiv:2506.15734v1 Announce Type: new
Abstract: As Vision-Language Models (VLMs) demonstrate increasing capabilities across real-world applications such as code generation and chatbot assistance, ensuring their safety has become paramount. Unlike traditional Large Language Models (LLMs), VLMs face unique vulnerabilities due to their multimodal nature, allowing adversaries to modify visual or textual inputs to bypass safety guardrails and trigger the generation of harmful content. Through systematic analysis of VLM behavior under attack, we identify a novel phenomenon termed “delayed safety awareness”. Specifically, we observe that safety-aligned VLMs may initially be compromised to produce harmful content, but eventually recognize the associated risks and attempt to self-correct. This pattern suggests that VLMs retain their underlying safety awareness but experience a temporal delay in their activation. Building on this insight, we hypothesize that VLMs’ safety awareness can be proactively reactivated through carefully designed prompts. To this end, we introduce “The Safety Reminder”, a soft prompt tuning approach that optimizes learnable prompt tokens, which are periodically injected during the text generation process to enhance safety awareness, effectively preventing harmful content generation. Additionally, our safety reminder only activates when harmful content is detected, leaving normal conversations unaffected and preserving the model’s performance on benign tasks. Through comprehensive evaluation across three established safety benchmarks and one adversarial attacks, we demonstrate that our approach significantly reduces attack success rates while maintaining model utility, offering a practical solution for deploying safer VLMs in real-world applications.

Source link

What's Hot

ChatGPT teen-safety measures to include age verification, OpenAI says

Daily Life of IBM’s Head of VC: Miles With Her Dogs and Meeting Startups

YouTube to use AI to help podcasters promote themselves with clips and Shorts

The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models

LTLCrit: A Temporal Logic-based LLM Critic for Safe and Efficient Embodied Agents

From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection

VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

Jennifer Packer and Marie Watt Win $250,000 Heinz Award

Sylvester Stallone Owns Works by Warhol, Condo, and Other Art Stars

LA Louver Gallery to Shutter Venice Gallery After 50 Years

Pritzker Family’s Hidden Art Trove Heads to Sotheby’s This Fall

ChatGPT teen-safety measures to include age verification, OpenAI says

Daily Life of IBM’s Head of VC: Miles With Her Dogs and Meeting Startups

YouTube to use AI to help podcasters promote themselves with clips and Shorts

What's Hot

The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models

Related Posts

Subscribe to Updates