저자: Dujun Nie, Xianda Guo, Yiqun Duan, Ruijun Zhang, Long Chen | 날짜: 2025-03-04 | URL: https://arxiv.org/abs/2503.02247 📄 PDF
라이선스: arXiv 비독점 라이선스
Fig. 2: The WMNav framework. After acquiring the RGB-D panoramic image and pose information at step t, the
Vision-Language Model을 기반으로 한 world model을 설계하여 Object Goal Navigation 작업에서 미래 상태를 예측하고 메모리를 통해 정책을 개선하는 WMNav 프레임워크를 제안한다. Curiosity Value Map이라는 온라인 유지 메모리 구조와 두 단계 행동 제안 전략으로 VLM의 hallucination을 완화하면서 탐색 효율성을 향상시킨다.
Fig. 2: The WMNav framework. After acquiring the RGB-D panoramic image and pose information at step t, the
Fig. 2: The WMNav framework. After acquiring the RGB-D panoramic image and pose information at step t, the
총평: 본 논문은 VLM을 world model로 활용하는 혁신적인 접근으로 zero-shot object navigation에서 새로운 방향을 제시하며, Curiosity Value Map 및 두 단계 행동 제안 전략이 효과적으로 탐색 효율성을 높인다. 체계적인 설계와 강력한 실험 결과로 embodied AI 분야에 중요한 기여를 한다.