Outdoor real-world environment
A photorealistic and continuous outdoor 3DGS campus environment for embodied VLM evaluation, supporting high-fidelity rendering and long-horizon interactive navigation.
A benchmark for Embodied Spatial Reasoning in large-scale, real-world outdoor 3D environments reconstructed by 3D Gaussian Splatting — unifying closed-loop interactive navigation with fine-grained diagnostic QA.
Embodied Spatial Reasoning (ESR) is central to deploying Vision-Language Models (VLMs) in real-world embodied tasks, yet remains inadequately evaluated. Existing benchmarks primarily assess spatial reasoning in static images or synthetic indoor environments, where semantic shortcuts often suffice, making it difficult to faithfully evaluate spatial reasoning capabilities in the wild. In light of this, we present Lost on Campus, a benchmark for evaluating ESR in large-scale real-world outdoor 3D environments reconstructed by 3D Gaussian Splatting.
Our benchmark introduces a unified reasoning-action evaluation framework that seamlessly integrates diagnostic QA for isolated reasoning with closed-loop interactive navigation for active reasoning under multimodal instructions. To enable fine-grained diagnosis, we systematically decompose ESR into six fundamental capabilities: action grounding, spatial foresight, metric awareness, goal-directed planning, Ego-allocentric Localization, and spatio-temporal consistency. Extensive experiments reveal that real-world outdoor environments impose substantially higher demands on spatial reasoning, that VLMs still struggle with fine-grained visual alignment — with spatial foresight and self-aware localization emerging as primary bottlenecks — and that improving multimodal interaction and long-term spatial reasoning is crucial for advancing embodied intelligence.
A photorealistic and continuous outdoor 3DGS campus environment for embodied VLM evaluation, supporting high-fidelity rendering and long-horizon interactive navigation.
We evaluate VLMs through closed-loop interactive navigation and diagnostic QA, combining reasoning in isolation with reasoning in action for fine-grained analysis of ESR capabilities.
We benchmark representative open- and closed-source VLMs, show outdoor ESR remains under-solved, and identify spatial foresight and Ego-allocentric Localization as the primary bottlenecks for embodied navigation.
Reconstructed from real-world large-scale captures with metric-accurate calibration.
Representative ground-truth and rendered view pairs across campus scenes.
Rendered along the original capture routes to show environmental scale, continuity, and visual fidelity.
General setting of 1,500 image-goal cases. Bold = best, underline = second-best.
| Model | SR@2.5 | SR@5 | OSR@2.5 | OSR@5 | SPL@2.5 | SPL@5 | DTG | Sim. | Prog. | Traj. | Calls | Actions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Proprietary Models | ||||||||||||
| Gemini-3-Flash | 25.73 | 49.60 | 36.13 | 67.73 | 17.29 | 31.16 | 10.11 | 0.699 | 0.27 | 39.14 | 21.7 | 66.2 |
| GPT-5-Mini | 11.80 | 30.60 | 21.47 | 49.60 | 9.54 | 24.34 | 13.81 | 0.615 | 0.22 | 35.45 | 24.4 | 67.5 |
| Open-source Models (~30B) | ||||||||||||
| Qwen3.5-27B | 9.93 | 34.07 | 11.33 | 36.47 | 9.05 | 31.91 | 9.48 | 0.699 | 0.51 | 14.49 | 28.7 | 44.6 |
| Qwen3-VL-32B-Instruct | 6.33 | 28.07 | 8.33 | 32.80 | 5.49 | 24.99 | 11.30 | 0.662 | 0.42 | 20.27 | 22.2 | 40.3 |
| Gemma-4-31B | 10.67 | 36.20 | 13.47 | 40.67 | 8.08 | 28.05 | 11.21 | 0.696 | 0.35 | 27.41 | 24.3 | 54.3 |
| Open-source Models (<10B) | ||||||||||||
| Qwen3.5-9B | 6.33 | 20.47 | 12.87 | 30.00 | 4.73 | 15.49 | 13.00 | 0.571 | 0.19 | 23.29 | 15.0 | 36.4 |
| Qwen3-VL-8B-Instruct | 10.67 | 25.53 | 17.27 | 39.13 | 8.58 | 20.78 | 13.81 | 0.606 | 0.20 | 27.12 | 40.9 | 60.2 |
| InternVL3.5-8B | 5.33 | 13.73 | 29.60 | 48.60 | 3.38 | 8.08 | 17.12 | 0.528 | 0.01 | 44.14 | 74.1 | 86.0 |
Models often approach the goal but fail to stop correctly. A persistent 20+ point gap between SR@2.5 and SR@5 — most extreme for InternVL3.5-8B (OSR@5 48.6% vs. SR@2.5 5.33%) — reveals that endpoint localization, goal verification, and self-aware stopping are dominant failure modes.
Per-capability accuracy (%). Bold = best, underline = second-best. The radar chart compares full capability profiles.
| Model | Action | Foresight | Metric | Planning | Localize | Spatio-temp. | Avg. |
|---|---|---|---|---|---|---|---|
| Baseline | |||||||
| Chance | 16.67 | 25.0 | 23.3 | 25.0 | 25.0 | 25.0 | 23.33 |
| Proprietary Models | |||||||
| Gemini-3-Flash | 75.16 | 61.9 | 62.2 | 75.4 | 41.7 | 62.4 | 63.13 |
| GPT-5-Mini | 58.53 | 53.3 | 50.6 | 61.4 | 40.0 | 55.2 | 53.17 |
| Open-source Models (~30B) | |||||||
| Qwen3.5-27B | 82.52 | 55.3 | 61.7 | 72.2 | 43.2 | 50.0 | 60.82 |
| Qwen3-VL-32B-Instruct | 65.25 | 52.2 | 44.7 | 65.7 | 29.1 | 53.8 | 51.79 |
| Gemma-4-31B | 76.01 | 58.2 | 59.0 | 72.5 | 44.4 | 54.3 | 60.74 |
| Open-source Models (<10B) | |||||||
| Qwen3.5-9B | 58.40 | 52.4 | 41.8 | 56.0 | 42.5 | 49.2 | 50.05 |
| Qwen3-VL-8B-Instruct | 52.36 | 30.5 | 38.0 | 53.0 | 43.9 | 48.5 | 44.38 |
| InternVL3.5-8B | 50.91 | 34.7 | 41.2 | 54.7 | 32.2 | 35.5 | 41.54 |
Ego-allocentric Localization and Spatial Foresight are the key bottlenecks. Both demand precise alignment between egocentric and allocentric reference frames — exactly what closed-loop navigation requires — indicating current VLMs lack the geometric grounding needed to "see themselves" in the world.
3DGS-based pose optimization applied after the agent stops (action sequence unchanged). Δ isolates the contribution of pose/localization error.
| Model | ΔSR@2.5 ↑ | ΔSR@5 ↑ | ΔDTG ↓ | ΔSim ↑ |
|---|---|---|---|---|
| Gemini-3-Flash | +10.00 | +2.33 | −0.53 | +0.026 |
| GPT-5-Mini | +6.40 | +3.80 | −0.34 | +0.012 |
| Qwen3.5-27B | +8.60 | +5.60 | −0.46 | +0.015 |
| Qwen3-VL-32B-Instruct | +7.14 | +5.66 | −0.41 | +0.011 |
| Gemma-4-31B | +10.00 | +5.13 | −0.47 | +0.014 |
| Qwen3.5-9B | +3.47 | +2.73 | −0.19 | +0.005 |
| Qwen3-VL-8B-Instruct | +5.33 | +2.74 | −0.30 | +0.014 |
| InternVL3.5-8B | +1.94 | +0.60 | −0.09 | +0.001 |
Larger gains for strong models (Gemini-3-Flash, Gemma-4-31B both +10.00 SR@2.5) indicate near-correct trajectories with terminal-pose misalignment, while sub-10B models carry deeper upstream reasoning failures.
255 controlled cases under the 5.0 m threshold. Image-Goal Nav. reports absolute SR / OSR (%); other rows report change (Δ) vs. Image-Goal Nav. positive change.
| Setting | Qwen3-VL-8B | Qwen3.5-9B | Qwen3.5-27B | Qwen3-VL-32B | ||||
|---|---|---|---|---|---|---|---|---|
| SR | OSR | SR | OSR | SR | OSR | SR | OSR | |
| Image-Goal Nav. (Abs.) | 4.69 | 9.38 | 0.39 | 0.78 | 3.91 | 4.69 | 5.14 | 6.17 |
| Δ w/ Language Instruction | −1.57 | −3.91 | +2.37 | +1.98 | +1.90 | +1.12 | +1.47 | +4.72 |
| Δ w/ Goal-Annotated BEV | +0.84 | −0.68 | +0.39 | 0.00 | +5.07 | +5.46 | +3.88 | +8.73 |
| Δ w/ Visual Route Map | −2.72 | −5.84 | +0.80 | +0.80 | +4.33 | +5.11 | +1.39 | +7.40 |
Structured visual-spatial cues beat language-only routes for larger VLMs. Qwen3.5-27B gains +5.07 (Goal-Annotated BEV) and +4.33 (Visual Route Map) SR, versus only +1.90 from Language Instruction — yet absolute SR stays below 10%, confirming the action-execution bottleneck.
Egocentric front/left/right views, goal images, and optional spatial instructions drive each decision step.
The agent observes its rendered view alongside the goal image, explores beyond the original camera path, and completes navigation through continuous 3DGS scenes.
Lost on Campus couples a photorealistic, continuously navigable 3DGS campus environment with closed-loop interactive navigation and diagnostic spatial QA, evaluating both spatial reasoning in action and fine-grained spatial understanding. Extensive experiments show that current VLMs remain limited in outdoor ESR, with spatial foresight and Ego-allocentric Localization emerging as key bottlenecks, and a persistent gap between spatial instruction and action execution. Faithful evaluation of ESR demands benchmarks that tightly couple perception, reasoning, and action within realistic, continuous environments.
@article{lostoncampus2026,
title = {Lost on Campus: Evaluating Embodied Spatial Reasoning of Vision-Language Models in the Wild},
author = {Zehan Zheng, Yanyuan Chen, Deming Li, Yutao Tang, Jianwen Xie, Alan Yuille, Jieneng Chen and Cheng Peng},
journal = {arXiv preprint},
year = {2026},
}