저자: Shuai Liu, Meng Cheng Lau | 날짜: 2025-09-23 | URL: https://arxiv.org/abs/2509.19023 📄 PDF
라이선스: arXiv 비독점 라이선스
Figure 1: Overview of the ROM-GRL framework. In Stage 1, a 4-DOF ROM policy is trained in Box2D: the policy
ROM-GRL은 모션캡처 데이터 없이 4-DOF reduced-order model로 생성한 gait template을 이용해 full-body humanoid 정책을 학습하는 2단계 강화학습 프레임워크이다. Adversarial discriminator를 통해 ROM의 5-dimensional gait feature 분포를 따르도록 유도하여 자연스러운 보행을 실현한다.
Figure 3 visualizes pelvis and foot trajectories for the ROM-GRL policy (blue) and the pure-reward baseline (orange),
Figure 2: Schematic of the planar ROM used to generate reference walking trajectories. The ROM consists of a central
총평: ROM-GRL은 reduced-order model을 creative하게 활용해 motion capture 의존성을 제거하면서 자연스럽고 안정적인 humanoid 보행을 달성하는 novel 프레임워크이다. 보상 설계와 모방 학습 간 간격을 효과적으로 줄였으나, 제한된 속도 범위와 실제 로봇 검증 부재가 일반화 가능성의 의문을 남긴다.