Known: VLA 모델은 embodiment별로 고품질 observation-action supervision을 필요로 하며, teleoperation은 controller-aligned label을 제공하지만 high-DoF humanoid에서는 확장이 느리고 비용이 크다. 반면 human egocentric video는 풍부한 bimanual manipulation 행동을 담고 있지만 실행 가능한 robot action label을 직접 제공하지 않는다.
Gap: human demonstration을 robot-executable supervision으로 변환하려면 embodiment alignment, observation-motion compatibility, action-interface alignment, joint-task consistency라는 네 가지 상호 결합된 요구사항을 동시에 만족해야 하는데, 기존 연구들은 perception label, hand trajectory, pretraining signal 수준에 머물러 controller-aligned 고자유도 joint supervision을 직접 산출하지 못한다.
Why: humanoid teleoperation의 데이터 수집 병목을 해소하면서도 실행 가능한 action supervision 품질을 유지할 수 있다면, 대규모 human video를 활용해 high-DoF humanoid VLA 학습의 데이터 확장성 문제를 근본적으로 완화할 수 있다.
Figure 1 Overview of Human-as-Humanoid and its effect on humanoid action-data generation. Conventional robot data
Human-as-Humanoid는 motion-recovery, robot-action-space, real-robot deployment 세 단계에서 변환 체인을 검증했으며, humanoid teleoperation 대비 4.8-7.2배의 raw demonstration-throughput 향상을 달성하고 target-task robot demonstration 없이 converted human label만으로 post-training된 policy가 실제 로봇 작업에 zero-shot 배포 가능함을 보였다.
How
Figure 4 Ego-exo data-collection setup. The egocentric cameras provide policy-aligned observations, while synchronized
총평: human video를 humanoid teleoperation의 대체·보완 데이터원으로 전환하는 실용적이고 체계적인 파이프라인을 제시하며, 실측 throughput 향상과 zero-shot 실로봇 배포로 실효성을 입증한 의미 있는 연구이다. 다만 검증 범위가 특정 embodiment와 제한된 task set에 집중되어 있어 더 폭넓은 일반화 검증이 뒤따라야 한다.