How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective

arXiv:2505.21505v1 Announce Type: cross
Abstract: Multilingual Alignment is an effective and representative paradigm to enhance LLMs’ multilingual capabilities, which transfers the capabilities from the high-resource languages to the low-resource languages. Meanwhile, some researches on language-specific neurons reveal that there are language-specific neurons that are selectively activated in LLMs when processing different languages. This provides a new perspective to analyze and understand LLMs’ mechanisms more specifically in multilingual scenarios. In this work, we propose a new finer-grained neuron identification algorithm, which detects language neurons~(including language-specific neurons and language-related neurons) and language-agnostic neurons. Furthermore, based on the distributional characteristics of different types of neurons, we divide the LLMs’ internal process for multilingual inference into four parts: (1) multilingual understanding, (2) shared semantic space reasoning, (3) multilingual output space transformation, and (4) vocabulary space outputting. Additionally, we systematically analyze the models before and after alignment with a focus on different types of neurons. We also analyze the phenomenon of ”Spontaneous Multilingual Alignment”. Overall, our work conducts a comprehensive investigation based on different types of neurons, providing empirical results and valuable insights for better understanding multilingual alignment and multilingual capabilities of LLMs.

Source link

What's Hot

LongCodeZip: Compress Long Context for Code Language Models – Takara TLDR

VIRTUE: Visual-Interactive Text-Image Universal Embedder – Takara TLDR

Vinod Khosla Slams ‘Tunnel Vision Creatives’ Attacking Sora As ‘AI Slop’

How does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective

LTLCrit: A Temporal Logic-based LLM Critic for Safe and Efficient Embodied Agents

From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection

VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

Former ARTnews Publisher Dies at 97

National Gallery of Art Closes as a Result of Government Shutdown

Almine Rech Closes London Gallery After More Than a Decade

Record Exec and Art Collector Gets Over 4 Years

LongCodeZip: Compress Long Context for Code Language Models – Takara TLDR

VIRTUE: Visual-Interactive Text-Image Universal Embedder – Takara TLDR

Vinod Khosla Slams ‘Tunnel Vision Creatives’ Attacking Sora As ‘AI Slop’

What's Hot

How does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective

Related Posts

Subscribe to Updates