Latent Action Diffusion for Cross-Embodiment Manipulation

저자: Erik Bauer, Elvis Nava, Robert K. Katzschmann | 날짜: 2025-06-17 | URL: https://arxiv.org/abs/2506.14608 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: arXiv 비독점 라이선스

Essence

Figure 1

Fig. 1: Overview of our approach. Left: We construct a semantically aligned latent action space by training modality-spe

로봇의 다양한 end-effector 간 action space 이질성을 극복하기 위해 contrastive learning으로 학습된 shared latent action space에서 diffusion policy를 학습하여 cross-embodiment 조작을 실현한다.

Motivation

Achievement

Figure 4

Fig. 4: Success rates for three different tasks comparing single-embodiment diffusion policies to cross-embodied latent

How

Figure 2

Fig. 2: The three-stage process for learning the cross-embodiment latent action space. Stage 1: Aligned end-effector (EE

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 3/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Cross-embodiment 로봇 학습의 action space 이질성 문제를 learned latent representation으로 우아하게 해결하고, contrastive learning과 diffusion policy를 조합하여 실제 성능 향상을 입증한 가치있는 연구이다. 다만 embodiment 다양성 범위 확대와 alignment 메커니즘의 더 깊은 분석이 후속 과제이다.

같이 보면 좋은 논문

기반 연구Diffusion Policy 논문은 action diffusion 기반 정책 학습의 기본 원리를 제안하며, Latent Action Diffusion의 이론적 바탕을 제공한다.
기반 연구CLAM: Continuous Latent Action Models는 latent action space에서 정책을 학습하는 기초적인 아이디어와 방법론을 Latent Action Diffusion에 제공합니다.
다른 접근Latent Action Diffusion 논문은 객체 상태가 아닌 latent action space를 통한 정책 일반화를 시도하며, 상태 표현의 차별점을 비교할 수 있다.
다른 접근Latent Action Diffusion for Cross-Embodiment Manipulation 논문은 다양한 로봇 embodiment간 스킬 전이에 diffusion approach를 활용하여 CrossFormer와 비교할 수 있다.
후속 연구Scaling Cross-Embodied Learning 논문은 cross-embodiment 정책 학습·확장성에 초점을 맞추어, Latent Action Diffusion의 핵심 아이디어를 대규모 실험·평가로 확장합니다.
후속 연구Latent Action Diffusion for Cross-Embodiment Manipulation 논문은 flow-matching 기반 action generation의 핵심 이론을 제공하여, NORA-1.5의 메커니즘 이해에 도움을 줍니다.
응용 사례1544는 다양한 시뮬레이션 환경에서 cross-embodiment policy의 실험 플랫폼을 제공하여, 1447의 latent action diffusion 방법의 실제적 평가에 도움을 준다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드