ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation
Vicky Feliren, A. Taufiq Asyhari, Muhamad Risqi U. Saputra
- Published
- Sep 15, 2026 — 17:42 UTC
Problem
The paper addresses a significant gap in uncertainty estimation for Vision-Language Navigation (VLN) models, particularly in the context of conformal prediction. The authors highlight that existing methods do not adequately account for the dependencies and variable lengths of VLN episodes, which can lead to unreliable uncertainty estimates. This work is presented as a preprint and has not yet undergone peer review.
Method
The core technical contribution is the Episode-Normalized Conformal Prediction (ENCP) framework. ENCP rescales the nonconformity score based on the policy's residual confidence, ensuring that each episode is calibrated with a single maximum score. The method utilizes data from the R2R and REVERIE datasets, incorporating four different VLN policies and three distinct nonconformity scores. Notably, the paper does not disclose specific details regarding the training compute used for the models.
Results
The ENCP method successfully met all reported empirical targets for step-coverage in a seen-to-unseen evaluation context. However, the available text does not report quantitative results comparing ENCP against specific baselines or benchmarks.
Limitations
The authors acknowledge that standard calibration techniques in conformal prediction do not provide coverage guarantees over the dependent and variable-length episodes typical in VLN tasks. This limitation suggests that while ENCP improves upon existing methods, there may still be challenges in achieving reliable uncertainty estimates across all scenarios.
Why it matters
The implications of this work are significant for downstream applications in VLN, as improved uncertainty estimation can enhance the robustness and reliability of navigation models. By addressing the limitations of traditional conformal prediction methods, ENCP may pave the way for more effective integration of uncertainty quantification in complex decision-making tasks involving vision and language.
By Callan Zhang · Sep 15, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
