Eigenspectrum Analysis Of Neural Networks Without Aspect Ratio Bias

arXiv:2506.06280v1 Announce Type: cross
Abstract: Diagnosing deep neural networks (DNNs) through the eigenspectrum of weight matrices has been an active area of research in recent years. At a high level, eigenspectrum analysis of DNNs involves measuring the heavytailness of the empirical spectral densities (ESD) of weight matrices. It provides insight into how well a model is trained and can guide decisions on assigning better layer-wise training hyperparameters. In this paper, we address a challenge associated with such eigenspectrum methods: the impact of the aspect ratio of weight matrices on estimated heavytailness metrics. We demonstrate that matrices of varying sizes (and aspect ratios) introduce a non-negligible bias in estimating heavytailness metrics, leading to inaccurate model diagnosis and layer-wise hyperparameter assignment. To overcome this challenge, we propose FARMS (Fixed-Aspect-Ratio Matrix Subsampling), a method that normalizes the weight matrices by subsampling submatrices with a fixed aspect ratio. Instead of measuring the heavytailness of the original ESD, we measure the average ESD of these subsampled submatrices. We show that measuring the heavytailness of these submatrices with the fixed aspect ratio can effectively mitigate the aspect ratio bias. We validate our approach across various optimization techniques and application domains that involve eigenspectrum analysis of weights, including image classification in computer vision (CV) models, scientific machine learning (SciML) model training, and large language model (LLM) pruning. Our results show that despite its simplicity, FARMS uniformly improves the accuracy of eigenspectrum analysis while enabling more effective layer-wise hyperparameter assignment in these application domains. In one of the LLM pruning experiments, FARMS reduces the perplexity of the LLaMA-7B model by 17.3% when compared with the state-of-the-art method.

Source link

What's Hot

OpenAI GPT-5 Possesses Doctorate-Level Abilities? Google DeepMind CEO: Nonsense_the_this_Market

Why the Dongfeng Yipai eπ007 Became a Best-Seller in the 150,000 Yuan New Energy Sedan Market_The_With

Innovaccer Revenue Surges Amidst $275M AI Funding Round

Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias

LTLCrit: A Temporal Logic-based LLM Critic for Safe and Efficient Embodied Agents

From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection

VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

Ohio Auction of Two Paintings Looted By Nazis Halted By Foundation

Lee Ufan Painting at Center of Bribery Investigation in Korea

Drought Reveals 40 Ancient Tombs in Northern Iraqi Reservoir

Artifacts Removed from Gaza Building Before Suspected Israeli Strike

OpenAI GPT-5 Possesses Doctorate-Level Abilities? Google DeepMind CEO: Nonsense_the_this_Market

Why the Dongfeng Yipai eπ007 Became a Best-Seller in the 150,000 Yuan New Energy Sedan Market_The_With

Innovaccer Revenue Surges Amidst $275M AI Funding Round

What's Hot

Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias

Related Posts

Subscribe to Updates