NeurIPS 2026 Evaluations & Datasets Track

Lost on Campus
Evaluating Embodied Spatial Reasoning of Vision-Language Models in the Wild

A benchmark for Embodied Spatial Reasoning in large-scale, real-world outdoor 3D environments reconstructed by 3D Gaussian Splatting — unifying closed-loop interactive navigation with fine-grained diagnostic QA.

Johns Hopkins University · University of Virginia · Stanford University · Lambda
50K+
Calibrated video frames
15
Outdoor 3DGS scenes
12.7K
m² average reconstruction area per scene
6
ESR capabilities
6,000
Diagnostic QA pairs
8
VLMs evaluated
Overview of the Lost on Campus benchmark
Overview of the Lost on Campus benchmark. The evaluation is built upon large-scale photorealistic outdoor 3DGS environments reconstructed from real-world campus captures. It leverages closed-loop interactive navigation under different spatial instructions, integrating spatial understanding with action execution, and uses navigation-grounded QA to provide a fine-grained probe of six fundamental ESR capabilities.
Abstract

Embodied Spatial Reasoning (ESR) is central to deploying Vision-Language Models (VLMs) in real-world embodied tasks, yet remains inadequately evaluated. Existing benchmarks primarily assess spatial reasoning in static images or synthetic indoor environments, where semantic shortcuts often suffice, making it difficult to faithfully evaluate spatial reasoning capabilities in the wild. In light of this, we present Lost on Campus, a benchmark for evaluating ESR in large-scale real-world outdoor 3D environments reconstructed by 3D Gaussian Splatting.

Our benchmark introduces a unified reasoning-action evaluation framework that seamlessly integrates diagnostic QA for isolated reasoning with closed-loop interactive navigation for active reasoning under multimodal instructions. To enable fine-grained diagnosis, we systematically decompose ESR into six fundamental capabilities: action grounding, spatial foresight, metric awareness, goal-directed planning, Ego-allocentric Localization, and spatio-temporal consistency. Extensive experiments reveal that real-world outdoor environments impose substantially higher demands on spatial reasoning, that VLMs still struggle with fine-grained visual alignment — with spatial foresight and self-aware localization emerging as primary bottlenecks — and that improving multimodal interaction and long-term spatial reasoning is crucial for advancing embodied intelligence.

Contributions

What this benchmark delivers

Outdoor real-world environment

A photorealistic and continuous outdoor 3DGS campus environment for embodied VLM evaluation, supporting high-fidelity rendering and long-horizon interactive navigation.

Unified reasoning-action framework

We evaluate VLMs through closed-loop interactive navigation and diagnostic QA, combining reasoning in isolation with reasoning in action for fine-grained analysis of ESR capabilities.

Empirical findings on outdoor ESR

We benchmark representative open- and closed-source VLMs, show outdoor ESR remains under-solved, and identify spatial foresight and Ego-allocentric Localization as the primary bottlenecks for embodied navigation.

The 3D Outdoor Environment

A photorealistic continuously navigable campus

Reconstructed from real-world large-scale captures with metric-accurate calibration.

Benchmark statistics and diagnostic QA task illustration
Benchmark statistics & task illustration. Left: coverage of buildings and intersections with large-scale navigable area, split into general (short-to-medium) and complex (long-distance) cases. Right: schematic of the six diagnostic QA tasks.
Reconstruction fidelity

Captured views and 3DGS renderings

Representative ground-truth and rendered view pairs across campus scenes.

Ground-truth and rendered campus view comparison 1 Ground-truth and rendered campus view comparison 2 Ground-truth and rendered campus view comparison 3 Ground-truth and rendered campus view comparison 4 Ground-truth and rendered campus view comparison 5 Ground-truth and rendered campus view comparison 6 Ground-truth and rendered campus view comparison 7 Ground-truth and rendered campus view comparison 8
Scene traversal

Continuous navigation through reconstructed scenes

Rendered along the original capture routes to show environmental scale, continuity, and visual fidelity.

Scene traversal 1 preview
94 rendered frames9.4 s
Scene traversal 2 preview
209 rendered frames20.9 s
Scene traversal 3 preview
198 rendered frames19.8 s
Benchmark results

Outdoor ESR remains under-solved

Closed-loop navigation

General setting of 1,500 image-goal cases. Bold = best, underline = second-best.

ModelSR@2.5SR@5OSR@2.5OSR@5 SPL@2.5SPL@5DTGSim.Prog. Traj.CallsActions
Proprietary Models
Gemini-3-Flash25.7349.6036.1367.7317.2931.1610.110.6990.2739.1421.766.2
GPT-5-Mini11.8030.6021.4749.609.5424.3413.810.6150.2235.4524.467.5
Open-source Models (~30B)
Qwen3.5-27B9.9334.0711.3336.479.0531.919.480.6990.5114.4928.744.6
Qwen3-VL-32B-Instruct6.3328.078.3332.805.4924.9911.300.6620.4220.2722.240.3
Gemma-4-31B10.6736.2013.4740.678.0828.0511.210.6960.3527.4124.354.3
Open-source Models (<10B)
Qwen3.5-9B6.3320.4712.8730.004.7315.4913.000.5710.1923.2915.036.4
Qwen3-VL-8B-Instruct10.6725.5317.2739.138.5820.7813.810.6060.2027.1240.960.2
InternVL3.5-8B5.3313.7329.6048.603.388.0817.120.5280.0144.1474.186.0

Models often approach the goal but fail to stop correctly. A persistent 20+ point gap between SR@2.5 and SR@5 — most extreme for InternVL3.5-8B (OSR@5 48.6% vs. SR@2.5 5.33%) — reveals that endpoint localization, goal verification, and self-aware stopping are dominant failure modes.

Diagnostic spatial QA

Per-capability accuracy (%). Bold = best, underline = second-best. The radar chart compares full capability profiles.

ModelActionForesightMetricPlanningLocalizeSpatio-temp.Avg.
Baseline
Chance16.6725.023.325.025.025.023.33
Proprietary Models
Gemini-3-Flash75.1661.962.275.441.762.463.13
GPT-5-Mini58.5353.350.661.440.055.253.17
Open-source Models (~30B)
Qwen3.5-27B82.5255.361.772.243.250.060.82
Qwen3-VL-32B-Instruct65.2552.244.765.729.153.851.79
Gemma-4-31B76.0158.259.072.544.454.360.74
Open-source Models (<10B)
Qwen3.5-9B58.4052.441.856.042.549.250.05
Qwen3-VL-8B-Instruct52.3630.538.053.043.948.544.38
InternVL3.5-8B50.9134.741.254.732.235.541.54
Radar chart of model capability profiles
Capability profiles. Models with similar averages can exhibit divergent strengths across the six ESR capabilities.

Ego-allocentric Localization and Spatial Foresight are the key bottlenecks. Both demand precise alignment between egocentric and allocentric reference frames — exactly what closed-loop navigation requires — indicating current VLMs lack the geometric grounding needed to "see themselves" in the world.

Additional experiments

Effect of post-hoc pose optimization

3DGS-based pose optimization applied after the agent stops (action sequence unchanged). Δ isolates the contribution of pose/localization error.

ModelΔSR@2.5 ↑ΔSR@5 ↑ΔDTG ↓ΔSim ↑
Gemini-3-Flash+10.00+2.33−0.53+0.026
GPT-5-Mini+6.40+3.80−0.34+0.012
Qwen3.5-27B+8.60+5.60−0.46+0.015
Qwen3-VL-32B-Instruct+7.14+5.66−0.41+0.011
Gemma-4-31B+10.00+5.13−0.47+0.014
Qwen3.5-9B+3.47+2.73−0.19+0.005
Qwen3-VL-8B-Instruct+5.33+2.74−0.30+0.014
InternVL3.5-8B+1.94+0.60−0.09+0.001

Larger gains for strong models (Gemini-3-Flash, Gemma-4-31B both +10.00 SR@2.5) indicate near-correct trajectories with terminal-pose misalignment, while sub-10B models carry deeper upstream reasoning failures.

Probing spatial reasoning in embodied interaction

255 controlled cases under the 5.0 m threshold. Image-Goal Nav. reports absolute SR / OSR (%); other rows report change (Δ) vs. Image-Goal Nav.   positive change.

Setting Qwen3-VL-8B Qwen3.5-9B Qwen3.5-27B Qwen3-VL-32B
SROSRSROSRSROSRSROSR
Image-Goal Nav. (Abs.) 4.699.380.390.783.914.695.146.17
Δ w/ Language Instruction −1.57−3.91 +2.37+1.98 +1.90+1.12 +1.47+4.72
Δ w/ Goal-Annotated BEV +0.84−0.68 +0.390.00 +5.07+5.46 +3.88+8.73
Δ w/ Visual Route Map −2.72−5.84 +0.80+0.80 +4.33+5.11 +1.39+7.40

Structured visual-spatial cues beat language-only routes for larger VLMs. Qwen3.5-27B gains +5.07 (Goal-Annotated BEV) and +4.33 (Visual Route Map) SR, versus only +1.90 from Language Instruction — yet absolute SR stays below 10%, confirming the action-execution bottleneck.

Inside the Loop

Agent Implementation Details

Egocentric front/left/right views, goal images, and optional spatial instructions drive each decision step.

Illustration of evaluation input and case visualization
Evaluation input & case visualization. "BEV Map Goal Instruction" denotes the Goal-Annotated Map and "BEV Map Route Instruction" denotes the Visual Route Map setting.
System prompt template for navigation
Unified system prompt template
User prompt template for navigation
Unified user prompt template
Gray text denotes optional setting-specific fields, conditionally included per evaluation setting.
Qualitative Results

Closed-loop navigation rollouts

The agent observes its rendered view alongside the goal image, explores beyond the original camera path, and completes navigation through continuous 3DGS scenes.

Closed-loop navigation rollout 1 preview
Final distance 2.11 m22 steps
Closed-loop navigation rollout 2 preview
Final distance 2.19 m54 steps
Closed-loop navigation rollout 3 preview
Final distance 2.07 m39 steps
Closed-loop navigation rollout 4 preview
Final distance 1.81 m18 steps
Closed-loop navigation rollout 5 preview
Final distance 1.52 m29 steps
Closed-loop navigation rollout 6 preview
Final distance 1.85 m47 steps

Diagnostic spatial QA

Examples across six ESR capabilities.


Diagnostic QA examples for Action Grounding, Spatial Foresight, Metric Awareness, Goal-directed Planning, Ego-allocentric Localization, and Spatio-temporal Consistency, with ground-truth answers and model predictions.
Diagnostic spatial QA examples. Each panel illustrates one ESR capability. Correct answers are highlighted in green.
Conclusion

Toward faithful embodied spatial reasoning

Lost on Campus couples a photorealistic, continuously navigable 3DGS campus environment with closed-loop interactive navigation and diagnostic spatial QA, evaluating both spatial reasoning in action and fine-grained spatial understanding. Extensive experiments show that current VLMs remain limited in outdoor ESR, with spatial foresight and Ego-allocentric Localization emerging as key bottlenecks, and a persistent gap between spatial instruction and action execution. Faithful evaluation of ESR demands benchmarks that tightly couple perception, reasoning, and action within realistic, continuous environments.

Citation

BibTeX

BibTeX
@article{lostoncampus2026,
  title     = {Lost on Campus: Evaluating Embodied Spatial Reasoning of Vision-Language Models in the Wild},
  author    = {Zehan Zheng, Yanyuan Chen, Deming Li, Yutao Tang, Jianwen Xie, Alan Yuille, Jieneng Chen and Cheng Peng},
  journal   = {arXiv preprint},
  year      = {2026},
}