FINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation

Khang H. Nguyen1,2,*,†, Hoang Pham Quang Nguyen1,*, Ha Phuong Nguyen1, Khanh Dinh Binh1, Xuan Ha Nguyen1, Vien Ngo1,3, Duy Ho Nguyen Minh4,5,6, Huan Nguyen1,‡, An T. Le1,3,7
1VinRobotics 2University of California, Los Angeles (UCLA) 3Center for AI Research, VinUniversity, Vietnam 4German Research Center for Artificial Intelligence (DFKI) 5University of Stuttgart 6Max Planck Research School for Intelligent Systems (IMPRS-IS) 7Intelligent Autonomous Systems, TU Darmstadt, Germany
*Equal contribution. †This work was done while the author was at VinRobotics. ‡Corresponding author: v.huannd13@vinrobotics.net

Project Video

Abstract

Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In particular, future observations reveal the instruction-relevant landmarks that the agent will encounter, including what they look like and how they are arranged in 3D.

We introduce FINE, a Future-Informed Navigation Encoding framework that extracts this latent supervision from existing demonstrations. FINE equips a VLN backbone with two complementary auxiliary representations. First, explicit landmark tokens follow the ordered landmarks specified by the instruction and are trained to predict the future landmark region in both semantic 2D patch-feature space and viewpoint-dependent 3D geometric feature space. Second, an implicit future token learns to distinguish the landmark state that is actually reached from plausible same-scene counterfactual futures generated by a video world model.

On R2R-CE and RxR-CE val-unseen, FINE improves InternVLA-N1 by 2.6 and 4.5 success-rate points, respectively, at full training data. More importantly, as demonstrations become limited, the benefit grows: at a 70% demonstration budget, FINE improves success rate by 6.8 points, recovering roughly one-third of the performance lost by reducing the training demonstrations.

Index Terms—Vision-Based Navigation, Deep Learning for Visual Perception, Representation Learning, Imitation Learning, AI-Enabled Robotics

Why FINE?

Adapting VLN policies to new environments is expensive: every new route and instruction requires an embodied demonstration. Most policies are trained through observation-to-action prediction alone, even though each demonstration already contains richer supervision — future observations reveal which landmark comes next, its appearance, and its 3D structure, with no extra annotation or interaction required.

Prior work either summarizes what the agent has already seen, or predicts future observations for use during inference. FINE takes a different view: the future should serve as a teacher during training, rather than an extra input or rollout mechanism at deployment.

Overview

Given the instruction and current observation, FINE builds ordered explicit landmark tokens and an implicit future-state token that condition the policy's action decoder — future supervision shapes training only, adding no extra computation at deployment.

FINE method overview.

Explicit Landmark Modeling

Navigation instructions typically describe an ordered sequence of landmarks — for example, pass the carton box, then head toward the blue table, and stop at the green plant. FINE represents these landmarks as ordered tokens that capture both identity and route position, grounds them in the current observation, and trains them to predict the semantic appearance and 3D geometric structure of the corresponding future landmark region.

Explicit landmark modeling component.

Implicit Navigation-State Modeling

Predicting landmarks teaches the model what should appear next, but not which future actually matches the demonstrated route. FINE addresses this with a future token trained to recognize the landmark state the demonstration actually reaches, while distinguishing it from plausible same-scene alternatives generated by a video world model.

Implicit navigation-state modeling component.

Main Results

Demonstration-budget comparison on R2R-CE val-unseen.

Budgets subsample training episodes; baseline and FINE consume the identical subset and the same number of optimizer updates. Within each baseline/FINE pair, the better value is in bold.

Budget Method NE↓ OS↑ SR↑ SPL↑
100% NaVILA 5.20 61.6 54.0 49.0
NaVILA + FINE 5.00 63.1 55.6 50.7
InternVLA-N1 4.83 63.3 58.2 54.0
InternVLA-N1 + FINE 4.29 65.7 60.8 55.2
70% NaVILA 6.76 46.9 36.2 29.5
NaVILA + FINE 6.19 50.6 41.7 35.5
InternVLA-N1 6.45 46.0 38.5 34.4
InternVLA-N1 + FINE 5.82 56.0 45.3 40.2
30% NaVILA 9.00 15.1 7.50 5.20
NaVILA + FINE 8.52 21.0 11.9 10.3
InternVLA-N1 7.79 13.2 9.20 8.30
InternVLA-N1 + FINE 7.46 21.3 14.4 13.5

Benchmark Results

VLN-CE R2R and RxR val-unseen results at 100% training data.

Prior results are from original papers and may differ in observation space and external data. Ours uses no external data; JanusVLN† uses ∼10.7M external image-action pairs. Pano.: panoramic RGB; Odo.: odometry; S.RGB: single-view RGB. Best in bold, second best underlined; —: not reported.

Method Observation R2R-CE Val-Unseen RxR-CE Val-Unseen
Pano. Odo. Depth S.RGB NE↓ OS↑ SR↑ SPL↑ NE↓ SR↑ SPL↑ nDTW↑
CMA ✓ ✓ ✓ 6.20 52.0 41.0 36.0 8.76 26.5 22.1 47.0
ETPNav ✓ ✓ ✓ 4.71 65.0 57.0 49.0 5.64 54.7 44.8 61.9
ScaleVLN ✓ ✓ ✓ 4.80 — 55.0 51.0 — — — —
NaVid ✓ 5.47 49.1 37.4 35.9 — — — —
UniNaVid ✓ 5.58 53.3 47.0 42.7 6.24 48.7 40.9 —
FutureNav ✓ 5.15 61.6 55.5 51.4 5.93 53.3 45.4 59.8
JanusVLN ✓ 5.17 58.0 52.8 49.2 6.46 51.4 44.3 59.1
JanusVLN† ✓ 4.78 65.2 60.5 56.8 6.06 56.2 47.5 62.1
NaVILA ✓ 5.22 62.5 54.0 49.0 6.77 49.3 44.0 58.8
StreamVLN ✓ 5.09 62.8 55.6 49.6 5.69 54.9 46.1 64.1
InternVLA-N1 ✓ 4.83 63.3 58.2 54.0 5.91 53.5 46.1 65.3
InternVLA-N1 + FINE (ours) ✓ 4.29 65.7 60.8 55.2 4.86 58.0 47.6 66.4

Qualitative Results

Simulation comparisons and real-world robot deployments.

Simulation

Each row: baseline | FINE   baseline | FINE (four videos).

InternVLA-N1 Baseline vs. + FINE
Failed InternVLA-N1
Video coming soon
Success + FINE
Video coming soon
Failed InternVLA-N1
Video coming soon
Success + FINE
Video coming soon
Failed InternVLA-N1
Video coming soon
Success + FINE
Video coming soon
Failed InternVLA-N1
Video coming soon
Success + FINE
Video coming soon
Failed InternVLA-N1
Video coming soon
Success + FINE
Video coming soon
Failed InternVLA-N1
Video coming soon
Success + FINE
Video coming soon
NaVILA Baseline vs. + FINE
Failed NaVILA
Video coming soon
Success + FINE
Video coming soon
Failed NaVILA
Video coming soon
Success + FINE
Video coming soon
Failed NaVILA
Video coming soon
Success + FINE
Video coming soon
Failed NaVILA
Video coming soon
Success + FINE
Video coming soon
Failed NaVILA
Video coming soon
Success + FINE
Video coming soon
Failed NaVILA
Video coming soon
Success + FINE
Video coming soon
Failed NaVILA
Video coming soon
Success + FINE
Video coming soon
Failed NaVILA
Video coming soon
Success + FINE
Video coming soon

Real-World Robot Deployment

FINE-enhanced navigation deployed on a physical robot. Each clip shows the onboard view with the natural-language instruction.

Instruction “Go along the glass wall, straight to the cartoon box. Then find the blue table, turn right and walk down the hallway. Stop at the green plant near the trash bin.”
Video coming soon
Instruction “Go to the blue table, then turn right. Stop in front of the yellow table.”
Video coming soon
Instruction “Go straight to the stack of cartoon box. Then turn left and walk down the glass wall. At the end, turn right, find the red fire distinguisher and stop. ”
Video coming soon
Instruction “Go straight pass the blue table. Turn right when you see the cartoon box. Go straight and stop at the green plant near the black glass door.”
Video coming soon
Instruction “Go to the carton box, turn left, and continue forward until you reach the green plant.”
Video coming soon
Instruction “Go to the carton box, then continue to the sofa, and head toward the blue table.”
Video coming soon

BibTeX

% BibTeX coming soon after the paper is released.