Abstract
Direct visual navigation policies can propose feasible-looking trajectories without explicitly evaluating their future consequences. LiteNWM adds a learned evaluator to a frozen navigation policy: shared visual encoding and joint prediction over four future horizons enable trajectory selection in latent space, without generating RGB videos. The robot executes a short action prefix before observing and replanning. We evaluate navigation accuracy and computational cost, and demonstrate image-goal navigation in unseen indoor and outdoor environments on a Unitree Go2 EDU.
Video
Latent World Modeling
Evaluate candidate trajectories through predicted visual features, without rendering future RGB frames.
Shared visual features → latent future prediction → trajectory selection. Select the figure to enlarge.
01 · Share the observation
Encode the visual history and goal once for all candidates.
02 · Predict latent futures
Jointly predict four horizons for 16 proposals and the incumbent.
03 · Select and execute
Score eligible challengers, execute a short action prefix, and observe again.
Real-World Navigation
On Unitree Go2 EDU, LiteNWM improves success rate from 43.3% to 83.3% over NoMaD [1] across three test scenes. The clips below illustrate representative runs; the table reports all 60 trials.
LiteNWM vs. NoMaD
Original-speed comparison · Play both runs together; speed is adjustable. Final-view insets remain visible for 5 seconds.
Quantitative results
| Scene | NoMaD SR | LiteNWM SR | NoMaD SPL | LiteNWM SPL |
|---|---|---|---|---|
| Open indoor | 70.0% | 90.0% | 0.642 | 0.890 |
| Cluttered indoor | 40.0% | 90.0% | 0.397 | 0.889 |
| Outdoor courtyard | 20.0% | 70.0% | 0.187 | 0.651 |
| Overall | 43.3% | 83.3% | 0.409 | 0.810 |
10 trials per method per scene · Higher SR and SPL are better. Paper · Table V ↗
Evaluation conditions and metrics
SR is the episode success rate. SPL measures success weighted by path efficiency, with failures assigned zero.
The policies use monocular RGB observations and an RGB goal. Success requires a goal-consistent view within a calibrated 1.0 m radius and 60° yaw tolerance, before collision, corrective intervention, or abnormal termination. The valid state must persist for 1 s or be followed by a normal stop.
NoMaD+NWM-S deployment probes are separate from these repeated-trial results. Protocol: Section IV-F and Table V ↗
Generalization Across Scenes
Qualitative navigation examples across unseen indoor and outdoor layouts, without scene-specific training or online model updates. Select a preview to enlarge.
Indoor · External view
Indoor · Robot view
Indoor · Robot view
Indoor · Robot view
Indoor · Robot view
Indoor · Robot view
Outdoor · External view
Outdoor · Robot view
Outdoor · Robot view
Robot views show sampled planning observations. External views are accelerated and include goal and robot-view insets; inset timing is approximate.
Long-Range Navigation
Successive visual subgoals guide longer routes; LiteNWM selects local trajectories along the way.
Indoor
Demo excerpt · 30× playback
Outdoor
Demo excerpt · 40× playback. Route map includes stitched segments; progress is illustrative.
Runtime and Computational Efficiency
On-Robot Planning Cost
In the NoMaD + NWM baseline, NoMaD [1] proposes candidate trajectories and Navigation World Models (NWM) [2] evaluates them using generated future RGB observations. These deployment examples use NWM-S, the smaller variant; both S and XL are evaluated in the RTX 5090 benchmark below.
Indoor NoMaD + NWM-S
150× playback · Click to play.
Outdoor NoMaD + NWM-S
100× playback · Click to play.
On-screen timers show elapsed run time. The clips use different playback speeds and are not synchronized comparisons.
Mean on-board planning time seconds per decision ↓
- NoMaD
- 0.651 s
- LiteNWM
- 2.228 s
- NoMaD + NWM-S
- 235.038 s
Internal planning time, excluding robot motion. NWM-S uses two exploratory runs and is excluded from the navigation success table.
On-board measurement scope
NWM-S uses H4/R1/DDIM-30 over two exploratory deployment runs, which are excluded from SR/SPL statistics. Latency values are aggregated over measurement runs; video timers describe elapsed episode time.
The RTX 5090 results below use a separate platform and measurement protocol; their speedup ratios do not describe on-board latency.
Computational Benchmark · RTX 5090
LiteNWM takes 0.691 s per state: 16.86× faster than NoMaD+NWM-S and 128× faster than NoMaD+NWM-XL in this benchmark.
Three-domain Macro ATE (lower is better). Bubble area: active parameters.
Table III · System performance on RTX 5090
| Method | Wall time (s) ↓ | GPU energy (Wh) ↓ | Peak VRAM (MiB) ↓ | Active params (M) |
|---|---|---|---|---|
| GNM | 0.076 | 0.0011 | 1,252 | 8.66 |
| ViNT | 0.088 | 0.0014 | 1,658 | 29.66 |
| NoMaD | 0.121 | 0.0021 | 791 | 19.05 |
| DINO-WM | 2.811 | 0.3560 | 3,520 | 41.30 |
| NWM-S | 11.635 | 1.5929 | 8,344 | 49.23 |
| NWM-XL | 88.449 | 14.070 | 13,776 | 1012 |
| NoMaD+NWM-S | 11.653 | 1.5929 | 8,344 | 154.40 |
| NoMaD+NWM-XL | 88.449 | 14.0695 | 13,776 | 1,117.13 |
| LiteNWM | 0.691 | 0.0277 | 2,459 | 349.04 |
Active parameters cover the whole workflow. Paper · Table III ↗
Compute measurements and chart guide
Wall time is the end-to-end computational cost per state, rather than navigation completion time. The pipeline covers proposal generation, future evaluation, scoring, refinement, and final selection. LiteNWM outer-process totals amortize initialization and warm-up over evaluated states.
Energy measures GPU consumption, not whole-robot energy. Peak VRAM is the maximum GPU memory usage. ATE is translational RMSE over the four future waypoints, as defined in the paper; lower is better. The bubble chart uses a logarithmic time axis.
Table III ↗ · Section IV-E ↗ · Resource measurement details ↗
References
- Ajay Sridhar, Dhruv Shah, Catherine Glossop, and Sergey Levine. NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration. ICRA 2024. Project ↗
- Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation World Models. CVPR 2025. Project ↗
See the paper for the complete bibliography and additional benchmark baselines.
