LiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the Wild

Linkai Liu1,2,†, Yuntian Zhang1,†, Zhenshan Bing1, Chen Chen3, Lingjuan Lyu3, Shangguang Wang4, Mengwei Xu4, Dongqi Cai1,*
1Nanjing University 2Imperial College London 3Sony AI 4Beijing University of Posts and Telecommunications

† Equal contribution   * Corresponding author

Indoor navigation 29.45× playback
Outdoor navigation 7.22× playback

LiteNWM selects navigation trajectories by predicting their consequences in latent space.

Model deployed locally on Unitree Go2 EDU   Watch the full video ↓

Abstract

Direct visual navigation policies can propose feasible-looking trajectories without explicitly evaluating their future consequences. LiteNWM adds a learned evaluator to a frozen navigation policy: shared visual encoding and joint prediction over four future horizons enable trajectory selection in latent space, without generating RGB videos. The robot executes a short action prefix before observing and replanning. We evaluate navigation accuracy and computational cost, and demonstrate image-goal navigation in unseen indoor and outdoor environments on a Unitree Go2 EDU.

Latent World Modeling

Evaluate candidate trajectories through predicted visual features, without rendering future RGB frames.

Shared visual features → latent future prediction → trajectory selection. Select the figure to enlarge.

01 · Share the observation

Encode the visual history and goal once for all candidates.

02 · Predict latent futures

Jointly predict four horizons for 16 proposals and the incumbent.

03 · Select and execute

Score eligible challengers, execute a short action prefix, and observe again.

Real-World Navigation

On Unitree Go2 EDU, LiteNWM improves success rate from 43.3% to 83.3% over NoMaD [1] across three test scenes. The clips below illustrate representative runs; the table reports all 60 trials.

LiteNWM vs. NoMaD

LiteNWM
NoMaD
0:00 / 0:22

Original-speed comparison · Play both runs together; speed is adjustable. Final-view insets remain visible for 5 seconds.

Quantitative results

SceneNoMaD SRLiteNWM SRNoMaD SPLLiteNWM SPL
Open indoor70.0%90.0%0.6420.890
Cluttered indoor40.0%90.0%0.3970.889
Outdoor courtyard20.0%70.0%0.1870.651
Overall43.3%83.3%0.4090.810

10 trials per method per scene · Higher SR and SPL are better. Paper · Table V ↗

Evaluation conditions and metrics

SR is the episode success rate. SPL measures success weighted by path efficiency, with failures assigned zero.

The policies use monocular RGB observations and an RGB goal. Success requires a goal-consistent view within a calibrated 1.0 m radius and 60° yaw tolerance, before collision, corrective intervention, or abnormal termination. The valid state must persist for 1 s or be followed by a normal stop.

NoMaD+NWM-S deployment probes are separate from these repeated-trial results. Protocol: Section IV-F and Table V ↗

Generalization Across Scenes

Qualitative navigation examples across unseen indoor and outdoor layouts, without scene-specific training or online model updates. Select a preview to enlarge.

Long-Range Navigation

Successive visual subgoals guide longer routes; LiteNWM selects local trajectories along the way.

Indoor

Demo excerpt · 30× playback

Outdoor

Demo excerpt · 40× playback. Route map includes stitched segments; progress is illustrative.

Runtime and Computational Efficiency

On-Robot Planning Cost

In the NoMaD + NWM baseline, NoMaD [1] proposes candidate trajectories and Navigation World Models (NWM) [2] evaluates them using generated future RGB observations. These deployment examples use NWM-S, the smaller variant; both S and XL are evaluated in the RTX 5090 benchmark below.

Indoor NoMaD + NWM-S

150× playback · Click to play.

Outdoor NoMaD + NWM-S

100× playback · Click to play.

On-screen timers show elapsed run time. The clips use different playback speeds and are not synchronized comparisons.

Mean on-board planning time seconds per decision ↓

NoMaD
0.651 s
LiteNWM
2.228 s
NoMaD + NWM-S
235.038 s

Internal planning time, excluding robot motion. NWM-S uses two exploratory runs and is excluded from the navigation success table.

On-board measurement scope

NWM-S uses H4/R1/DDIM-30 over two exploratory deployment runs, which are excluded from SR/SPL statistics. Latency values are aggregated over measurement runs; video timers describe elapsed episode time.

The RTX 5090 results below use a separate platform and measurement protocol; their speedup ratios do not describe on-board latency.

Computational Benchmark · RTX 5090

LiteNWM takes 0.691 s per state: 16.86× faster than NoMaD+NWM-S and 128× faster than NoMaD+NWM-XL in this benchmark.

Three-domain Macro ATE (lower is better). Bubble area: active parameters.

Table III · System performance on RTX 5090

MethodWall time (s) ↓GPU energy (Wh) ↓Peak VRAM (MiB) ↓Active params (M)
GNM0.0760.00111,2528.66
ViNT0.0880.00141,65829.66
NoMaD0.1210.002179119.05
DINO-WM2.8110.35603,52041.30
NWM-S11.6351.59298,34449.23
NWM-XL88.44914.07013,7761012
NoMaD+NWM-S11.6531.59298,344154.40
NoMaD+NWM-XL88.44914.069513,7761,117.13
LiteNWM0.6910.02772,459349.04

Active parameters cover the whole workflow. Paper · Table III ↗

Compute measurements and chart guide

Wall time is the end-to-end computational cost per state, rather than navigation completion time. The pipeline covers proposal generation, future evaluation, scoring, refinement, and final selection. LiteNWM outer-process totals amortize initialization and warm-up over evaluated states.

Energy measures GPU consumption, not whole-robot energy. Peak VRAM is the maximum GPU memory usage. ATE is translational RMSE over the four future waypoints, as defined in the paper; lower is better. The bubble chart uses a logarithmic time axis.

Table III ↗ · Section IV-E ↗ · Resource measurement details ↗

References

  1. Ajay Sridhar, Dhruv Shah, Catherine Glossop, and Sergey Levine. NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration. ICRA 2024. Project ↗
  2. Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation World Models. CVPR 2025. Project ↗

See the paper for the complete bibliography and additional benchmark baselines.