Adapting vision-language navigation (VLN) policies to new environments is expensive because every
additional route and instruction requires an embodied demonstration. Yet standard observation-to-action
training uses only a small fraction of the information already contained in each trajectory. In
particular, future observations reveal the instruction-relevant landmarks that the agent will encounter,
including what they look like and how they are arranged in 3D.
We introduce FINE, a Future-Informed
Navigation Encoding framework that extracts this latent supervision from
existing demonstrations. FINE equips a VLN backbone with two complementary auxiliary representations.
First, explicit landmark tokens follow the ordered landmarks specified by the instruction and are
trained to predict the future landmark region in both semantic 2D patch-feature space and
viewpoint-dependent 3D geometric feature space. Second, an implicit future token learns to
distinguish the landmark state that is actually reached from plausible same-scene counterfactual futures
generated by a video world model.
On R2R-CE and RxR-CE val-unseen, FINE improves InternVLA-N1 by 2.6 and 4.5 success-rate points,
respectively, at full training data. More importantly, as demonstrations become limited, the benefit
grows: at a 70% demonstration budget, FINE improves success rate by 6.8 points, recovering roughly
one-third of the performance lost by reducing the training demonstrations.
Index Terms—Vision-Based Navigation, Deep Learning for Visual Perception, Representation
Learning, Imitation Learning, AI-Enabled Robotics