Motus: A Unified Latent Action World Model

저자: Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, Hongyan Zhao, Hanyu Liu, Zhizhong Su, Lei Ma, Hang Su, Jun Zhu | 날짜: 2025-12-15 | URL: https://arxiv.org/abs/2512.13030 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: arXiv 비독점 라이선스

Essence

Figure 1

Figure 1. Motus Architecture. Here, at . . . at+k are actions, zt . . . zt+k are latent actions, and τv and τa are the r

Motus는 vision-language-action 모델, world 모델, inverse dynamics 모델, video generation 모델을 unified latent action world model로 통합하는 embodied agent 프레임워크이며, Mixture-of-Transformer 아키텍처와 optical flow 기반 latent action을 통해 대규모 이질적 데이터 학습을 가능하게 한다.

Motivation

Achievement

Figure 1

Figure 1. Motus Architecture. Here, at . . . at+k are actions, zt . . . zt+k are latent actions, and τv and τa are the r

How

Figure 1

Figure 1. Motus Architecture. Here, at . . . at+k are actions, zt . . . zt+k are latent actions, and τv and τa are the r

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Motus는 분산된 embodied agent 아키텍처를 unified model로 통합하면서 optical flow 기반 latent action과 체계적인 multi-stage 학습으로 대규모 이질적 데이터 활용을 가능하게 한 혁신적 연구이며, 강력한 실험 성과와 함께 embodied AI의 통합 모델링에 대한 새로운 패러다임을 제시한다.

같이 보면 좋은 논문

기반 연구DreamDojo는 대규모 휴먼 행동 world modeling 측면에서 Motus와 같은 VLA-World Model 통합의 이론적 기초를 제공한다.
기반 연구Motus는 latent action world model을 통해 DemoDiffusion의 kinematic retargeting 이후의 동작다양성 및 생성 능력을 한층 더 확장합니다.
다른 접근1481은 통합 latent action world model로 다양한 비디오 기반 행동 학습을 실현하며, 1448과 마찬가지로 라벨 없는 데이터에서 로봇 행동을 추출하는 접근을 제시한다.
다른 접근Diffusion-VLA 논문은 diffusion 기반 latent action policy로, Motus의 mixture-of-transformer 기반 world model 접근과 대비됩니다.
다른 접근Unified Video Action Model도 시각-행동-동영상 예측을 통합하는 접근으로 Motus의 unified latent action world model 구조와 유사점을 보여준다.
다른 접근Motus는 로코-매니퓰레이션을 위한 unified latent action world model을 제안하여, WholeBodyVLA와 비교되는 대안적 잠재공간 제어 접근법을 보여준다.
다른 접근Motus는 short/long-term memory 구성을 넘어 video-action joint prediction을 활용하여 장기간 작업 처리 방식의 또 다른 예시를 제공한다.
후속 연구Motus는 unified latent action world model을 제안하여, EnerVerse의 4D Gaussian Splatting 기반 world 모델 구성에 개념적 기초를 제공합니다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드