Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments

저자: Xiaopeng Lin, Ruoqi Yang, Shijie Lian, Zhaolong Shen, Bin Yu, Changti Wu, Haibao Liu, Yuxiang Zhang, Hong Li, Qiyuan Su, Haochen Liu, Xuguo He, Yukun Shi, Cong Huang, Zhirui Zhang, Bojun Cheng, Kai Chen | 날짜: 2026-06-30 | DOI: 10.48550/arXiv.2606.32009


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

Essence

Figure 1

Figure 1 Overview of Human-as-Humanoid and its effect on humanoid action-data generation. Conventional robot data

Human-as-Humanoid은 synchronized ego-exo human videos를 human-aligned 60-DoF upper-body humanoid인 PrimeU의 controller-aligned action label로 변환하여, robot teleoperation 없이 humanoid VLA policy를 학습·배포하는 human-to-humanoid supervision 프레임워크이다.

Motivation

Achievement

Figure 1

Figure 1 Overview of Human-as-Humanoid and its effect on humanoid action-data generation. Conventional robot data

Human-as-Humanoid는 motion-recovery, robot-action-space, real-robot deployment 세 단계에서 변환 체인을 검증했으며, humanoid teleoperation 대비 4.8-7.2배의 raw demonstration-throughput 향상을 달성하고 target-task robot demonstration 없이 converted human label만으로 post-training된 policy가 실제 로봇 작업에 zero-shot 배포 가능함을 보였다.

How

Figure 4

Figure 4 Ego-exo data-collection setup. The egocentric cameras provide policy-aligned observations, while synchronized

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: human video를 humanoid teleoperation의 대체·보완 데이터원으로 전환하는 실용적이고 체계적인 파이프라인을 제시하며, 실측 throughput 향상과 zero-shot 실로봇 배포로 실효성을 입증한 의미 있는 연구이다. 다만 검증 범위가 특정 embodiment와 제한된 task set에 집중되어 있어 더 폭넓은 일반화 검증이 뒤따라야 한다.

같이 보면 좋은 논문

기반 연구dexterous visual-tactile-action 데이터셋이 zero-shot humanoid 학습의 기반이 된다.
후속 연구humanoid manipulation 데이터셋 구축이 zero-shot 학습 기반이 된다.
다른 접근humanoid 학습을 위한 서로 다른 sim-to-real 전략을 제시한다.
후속 연구인간 영상 기반 humanoid 행동 학습이라는 공통 방법론적 기반을 공유한다.
응용 사례synchronized human video 기반 학습이 loco-manipulation 궤적 생성에 적용된다.
후속 연구humanoid RL 기반 동작 학습이라는 공통 기법을 공유한다.
후속 연구embodiment에 무관한 고수준 인지 표현 학습이라는 공통된 기반 개념을 공유한다.
후속 연구인간 영상 기반 학습 방법이 heavy-payload teleoperation으로 확장된다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드