SPhyR: Spatial-Physical Reasoning Benchmark On Material Distribution

arXiv:2505.16048v1 Announce Type: new
Abstract: We introduce a novel dataset designed to benchmark the physical and spatial reasoning capabilities of Large Language Models (LLM) based on topology optimization, a method for computing optimal material distributions within a design space under prescribed loads and supports. In this dataset, LLMs are provided with conditions such as 2D boundary, applied forces and supports, and must reason about the resulting optimal material distribution. The dataset includes a variety of tasks, ranging from filling in masked regions within partial structures to predicting complete material distributions. Solving these tasks requires understanding the flow of forces and the required material distribution under given constraints, without access to simulation tools or explicit physical models, challenging models to reason about structural stability and spatial organization. Our dataset targets the evaluation of spatial and physical reasoning abilities in 2D settings, offering a complementary perspective to traditional language and logic benchmarks.

Source link

What's Hot

New MIT Tech Sees Underwater As if the Water Weren’t There

AI fuels false claims after Charlie Kirk’s death, CBS News analysis reveals

Google is a ‘bad actor’ says People CEO, accusing the company of stealing content

SPhyR: Spatial-Physical Reasoning Benchmark on Material Distribution

LTLCrit: A Temporal Logic-based LLM Critic for Safe and Efficient Embodied Agents

From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection

VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

Ohio Auction of Two Paintings Looted By Nazis Halted By Foundation

Lee Ufan Painting at Center of Bribery Investigation in Korea

Nicholas Galanin Pulls Out of Smithsonian Event, Claiming Censorship

Two More Staffers Fired from Kennedy Center after Trump Takeover

New MIT Tech Sees Underwater As if the Water Weren’t There

AI fuels false claims after Charlie Kirk’s death, CBS News analysis reveals

Google is a ‘bad actor’ says People CEO, accusing the company of stealing content

What's Hot

SPhyR: Spatial-Physical Reasoning Benchmark on Material Distribution

Related Posts

Subscribe to Updates