From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers

arXiv:2506.17052v1 Announce Type: cross
Abstract: Transformers have achieved state-of-the-art performance across language and vision tasks. This success drives the imperative to interpret their internal mechanisms with the dual goals of enhancing performance and improving behavioral control. Attribution methods help advance interpretability by assigning model outputs associated with a target concept to specific model components. Current attribution research primarily studies multi-layer perceptron neurons and addresses relatively simple concepts such as factual associations (e.g., Paris is located in France). This focus tends to overlook the impact of the attention mechanism and lacks a unified approach for analyzing more complex concepts. To fill these gaps, we introduce Scalable Attention Module Discovery (SAMD), a concept-agnostic method for mapping arbitrary, complex concepts to specific attention heads of general transformer models. We accomplish this by representing each concept as a vector, calculating its cosine similarity with each attention head, and selecting the TopK-scoring heads to construct the concept-associated attention module. We then propose Scalar Attention Module Intervention (SAMI), a simple strategy to diminish or amplify the effects of a concept by adjusting the attention module using only a single scalar parameter. Empirically, we demonstrate SAMD on concepts of varying complexity, and visualize the locations of their corresponding modules. Our results demonstrate that module locations remain stable before and after LLM post-training, and confirm prior work on the mechanics of LLM multilingualism. Through SAMI, we facilitate jailbreaking on HarmBench (+72.7%) by diminishing “safety” and improve performance on the GSM8K benchmark (+1.6%) by amplifying “reasoning”. Lastly, we highlight the domain-agnostic nature of our approach by suppressing the image classification accuracy of vision transformers on ImageNet.

Source link

What's Hot

Thinking Machines Lab’s $2B Seed Round Is Biggest By A Long Shot

Musk’s attempts to politicize his Grok AI are bad for users and enterprises — here’s why

Leak reveals Grok might soon edit your spreadsheets

From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers

Linear-Time Primitives for Algorithm Development in Graphical Causal Inference

Circuit Analysis in LLMs for Propositional Logical Reasoning

[2506.11618] Convergent Linear Representations of Emergent Misalignment

Publicity Wizard Jalila Singerff On The Vital PR Rules For 2025

Romania Wins ‘Hold’ on El Greco, Arnaldo Pomodoro Dead, and More

Empire Of The Sun’s Luke Steele On Loss, Grief, Al Green And More

5 Standout Exhibitions To See In Venice During The Architecture Biennale

Thinking Machines Lab’s $2B Seed Round Is Biggest By A Long Shot

Musk’s attempts to politicize his Grok AI are bad for users and enterprises — here’s why

Leak reveals Grok might soon edit your spreadsheets

What's Hot

From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers

Related Posts

Subscribe to Updates