Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation
Date:
Large language models act as planners in tasks with inherent spatial structure, yet stay brittle at sequential spatial reasoning. Rather than asking whether they fail at navigation, the poster asks where in the spatial-cognition pipeline they get lost — decomposing maze navigation into Fine (local passability), Meso (junction topology) and Macro (goal orientation) levels across 1,050 mazes. First errors concentrate on Meso junction choices (59%) and Fine perception (39%), with global heading almost never at fault (1%).
Download the poster (PDF) · Project page and code · Paper
