저자: Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang, Zhiyang Liu, Ran Chen, Liqiang Nie | 날짜: 2025-08-18 | URL: https://arxiv.org/abs/2508.13073 📄 PDF
라이선스: arXiv 비독점 라이선스
Fig. 2: Outline of the organization of our comprehensive survey (top) and a chronological timeline of notable developmen
대규모 Vision-Language Model(VLM)을 기반으로 한 Vision-Language-Action(VLA) 모델들을 로봇 매니퓰레이션에 적용하는 연구의 첫 번째 체계적 설문조사로, Monolithic 모델과 Hierarchical 모델이라는 두 가지 주요 아키텍처 패러다임을 제시한다.
Fig. 3: Comparison of the two principal categories of large VLM-based VLA models. Monolithic models (Sec. 3) integrate
Fig. 2: Outline of the organization of our comprehensive survey (top) and a chronological timeline of notable developmen
총평: 본 설문조사는 빠르게 성장하는 VLM 기반 VLA 분야의 첫 번째 체계적 종합으로, 명확한 정의, 일관된 분류체계, 그리고 포괄적 분석을 통해 학계의 연구 단편화를 해소하고 향후 발전 방향을 제시하는 의의가 크다. 정기적 업데이트 계획도 분야의 빠른 진전을 반영하는 강점이다.