<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>physical-ai — Paper Curation</title>
  <link href="https://paper-curation.jehyunlee.dev/physical-ai/" rel="alternate" type="text/html"/>
  <link href="https://paper-curation.jehyunlee.dev/physical-ai/feed.xml" rel="self" type="application/atom+xml"/>
  <id>https://paper-curation.jehyunlee.dev/physical-ai/</id>
  <updated>2026-03-01T00:00:00Z</updated>
  <author><name>Jehyun Lee</name></author>
  <generator>paper-curation build_rss.py</generator>
  <entry>
    <title>MEM: Multi-Scale Embodied Memory for Vision Language Action Models</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1474_MEM_Multi-Scale_Embodied_Memory_for_Vision_Language_Action_M/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1474_MEM_Multi-Scale_Embodied_Memory_for_Vision_Language_Action_M/</id>
    <updated>2026-03-01T00:00:00Z</updated>
    <author><name>Marcel Torne</name></author>
    <author><name>Karl Pertsch</name></author>
    <author><name>Homer Walke</name></author>
    <author><name>Kyle Vedder</name></author>
    <author><name>Suraj Nair</name></author>
    <category term="LLM-Augmented Embodied Agent Frameworks"/>
    <summary type="text">로봇의 장시간 작업을 위해 비디오 기반 단기 메모리와 텍스트 기반 장기 메모리를 결합한 Multi-Scale Embodied Memory (MEM)을 제안하여, 15분 이상의 복잡한 조작 작업을 수행할 수 있는 Vision Language Action 모델을 구현했다.</summary>
  </entry>
  <entry>
    <title>EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1376_EgoScale_Scaling_Dexterous_Manipulation_with_Diverse_Egocent/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1376_EgoScale_Scaling_Dexterous_Manipulation_with_Diverse_Egocent/</id>
    <updated>2026-02-01T00:00:00Z</updated>
    <author><name>Ruijie Zheng</name></author>
    <author><name>Dantong Niu</name></author>
    <author><name>Yuqi Xie</name></author>
    <author><name>Jing Wang</name></author>
    <author><name>Mengda Xu</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">20,854시간의 대규모 이고센트릭 인간 비디오 데이터로 VLA 모델을 사전학습한 후 소량의 정렬된 인간-로봇 중간학습 데이터로 미세조정하여 22-DoF 손가락 조작 로봇에서 54% 성공률 향상을 달성했다.</summary>
  </entry>
  <entry>
    <title>DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1355_DreamDojo_A_Generalist_Robot_World_Model_from_Large-Scale_Hu/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1355_DreamDojo_A_Generalist_Robot_World_Model_from_Large-Scale_Hu/</id>
    <updated>2026-02-01T00:00:00Z</updated>
    <author><name>Shenyuan Gao</name></author>
    <author><name>William Liang</name></author>
    <author><name>Kaiyuan Zheng</name></author>
    <author><name>Ayaan Malik</name></author>
    <author><name>Seonghyeon Ye</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">44k시간의 대규모 인간 동영상으로부터 연속 잠재 행동(continuous latent actions)을 통일된 프록시로 사용하여 학습한 DreamDojo는 로봇의 손재주 제어와 물리 이해를 갖춘 기초 세계 모델로, 실시간 텔레오퍼레이션과 모델 기반 계획을 가능하게 한다.</summary>
  </entry>
  <entry>
    <title>PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1517_PointWorld_Scaling_3D_World_Models_for_In-The-Wild_Robotic_M/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1517_PointWorld_Scaling_3D_World_Models_for_In-The-Wild_Robotic_M/</id>
    <updated>2026-01-01T00:00:00Z</updated>
    <author><name>Wenlong Huang</name></author>
    <author><name>Yu-Wei Chao</name></author>
    <author><name>Arsalan Mousavian</name></author>
    <author><name>Ming-Yu Liu</name></author>
    <author><name>Dieter Fox</name></author>
    <category term="3D Simulation and Robot Manipulation"/>
    <summary type="text">PointWorld는 RGB-D 입력과 로봇 동작을 3D point flow로 통일하여 표현하고, 이를 통해 전체 장면의 3D 포인트 변위를 예측하는 대규모 사전학습 3D 월드 모델이다. 단일 체크포인트로 실제 로봇이 다양한 조작 작업을 수행할 수 있게 한다.</summary>
  </entry>
  <entry>
    <title>InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1437_InternVLA-A1_Unifying_Understanding_Generation_and_Action_fo/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1437_InternVLA-A1_Unifying_Understanding_Generation_and_Action_fo/</id>
    <updated>2026-01-01T00:00:00Z</updated>
    <author><name>Junhao Cai</name></author>
    <author><name>Zetao Cai</name></author>
    <author><name>Jiafei Cao</name></author>
    <author><name>Yilun Chen</name></author>
    <author><name>Zeyu He</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">InternVLA-A1은 Mixture-of-Transformers 아키텍처를 통해 의미 이해, 시각적 예측, 행동 실행을 통합하여 로봇 조작 성능을 향상시키는 Vision-Language-Action 모델이다. 실세계 로봇 데이터, 합성 시뮬레이션 데이터, 인간 비디오를 포함한 692M 프레임의 이질적 데이터로 사전학습되어 동적 조작 작업에서 26.7% 성능 향상을 달성한다.</summary>
  </entry>
  <entry>
    <title>DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1374_DynamicVLA_A_Vision-Language-Action_Model_for_Dynamic_Object/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1374_DynamicVLA_A_Vision-Language-Action_Model_for_Dynamic_Object/</id>
    <updated>2026-01-01T00:00:00Z</updated>
    <author><name>Haozhe Xie</name></author>
    <author><name>Beichen Wen</name></author>
    <author><name>Jiarui Zheng</name></author>
    <author><name>Zhaoxi Chen</name></author>
    <author><name>Fangzhou Hong</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">DynamicVLA는 동적 객체 조작을 위한 compact 0.4B VLA 모델로, Continuous Inference와 Latent-aware Action Streaming을 통해 지각-실행 간의 지연을 제거하고 실시간 폐루프 제어를 가능하게 한다.</summary>
  </entry>
  <entry>
    <title>Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1318_Being-H05_Scaling_Human-Centric_Robot_Learning_for_Cross-Emb/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1318_Being-H05_Scaling_Human-Centric_Robot_Learning_for_Cross-Emb/</id>
    <updated>2026-01-01T00:00:00Z</updated>
    <author><name>Hao Luo</name></author>
    <author><name>Ye Wang</name></author>
    <author><name>Wanpeng Zhang</name></author>
    <author><name>Sipeng Zheng</name></author>
    <author><name>Ziheng Xi</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">Being-H0.5는 인간 중심 학습 패러다임과 통합 액션 공간을 활용하여 다양한 로봇 플랫폼 간 일반화를 가능하게 하는 기초 Vision-Language-Action 모델이다. 35,000시간 이상의 멀티모달 데이터로 구성된 UniHand-2.0을 통해 30개의 로봇 플랫폼에서 강력한 cross-embodiment 성능을 달성한다.</summary>
  </entry>
  <entry>
    <title>A Pragmatic VLA Foundation Model</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1296_A_Pragmatic_VLA_Foundation_Model/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1296_A_Pragmatic_VLA_Foundation_Model/</id>
    <updated>2026-01-01T00:00:00Z</updated>
    <author><name>Wei Wu</name></author>
    <author><name>Fan Lu</name></author>
    <author><name>Yunnan Wang</name></author>
    <author><name>Shuai Yang</name></author>
    <author><name>Shi Liu</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">LingBot-VLA는 약 20,000시간의 실제 로봇 데이터로 학습한 Vision-Language-Action 기초 모델로, 효율적인 학습과 다중 플랫폼 일반화 능력을 갖춘다.</summary>
  </entry>
  <entry>
    <title>WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1628_WholeBodyVLA_Towards_Unified_Latent_VLA_for_Whole-Body_Loco-/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1628_WholeBodyVLA_Towards_Unified_Latent_VLA_for_Whole-Body_Loco-/</id>
    <updated>2025-12-01T00:00:00Z</updated>
    <author><name>Haoran Jiang</name></author>
    <author><name>Jin Chen</name></author>
    <author><name>Qingwen Bu</name></author>
    <author><name>Li Chen</name></author>
    <author><name>Modi Shi</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">WholeBodyVLA는 Vision-Language-Action 프레임워크로 humanoid 로봇의 대규모 공간에서 end-to-end 전신 조작-이동(loco-manipulation) 제어를 가능하게 한다. Unified latent learning으로 저비용 영상에서 학습하고 LMO RL policy로 정확한 이동 실행을 보장한다.</summary>
  </entry>
  <entry>
    <title>Motus: A Unified Latent Action World Model</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1481_Motus_A_Unified_Latent_Action_World_Model/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1481_Motus_A_Unified_Latent_Action_World_Model/</id>
    <updated>2025-12-01T00:00:00Z</updated>
    <author><name>Hongzhe Bi</name></author>
    <author><name>Hengkai Tan</name></author>
    <author><name>Shenghao Xie</name></author>
    <author><name>Zeyuan Wang</name></author>
    <author><name>Shuhe Huang</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">Motus는 vision-language-action 모델, world 모델, inverse dynamics 모델, video generation 모델을 unified latent action world model로 통합하는 embodied agent 프레임워크이며, Mixture-of-Transformer 아키텍처와 optical flow 기반 latent action을 통해 대규모 이질적 데이터 학습을 가능하게 한다.</summary>
  </entry>
  <entry>
    <title>HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1424_HiMoE-VLA_Hierarchical_Mixture-of-Experts_for_Generalist_Vis/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1424_HiMoE-VLA_Hierarchical_Mixture-of-Experts_for_Generalist_Vis/</id>
    <updated>2025-12-01T00:00:00Z</updated>
    <author><name>Zhiying Du</name></author>
    <author><name>Bei Liu</name></author>
    <author><name>Yaobo Liang</name></author>
    <author><name>Yichao Shen</name></author>
    <author><name>Haidong Cao</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">HiMoE-VLA는 로봇 데이터의 이질성(action space, embodiment, sensor configuration 등)을 명시적으로 처리하기 위해 계층적 Mixture-of-Experts 아키텍처를 제안하는 Vision-Language-Action 프레임워크이다.</summary>
  </entry>
  <entry>
    <title>Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1414_Ground_Slow_Move_Fast_A_Dual-System_Foundation_Model_for_Gen/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1414_Ground_Slow_Move_Fast_A_Dual-System_Foundation_Model_for_Gen/</id>
    <updated>2025-12-01T00:00:00Z</updated>
    <author><name>Meng Wei</name></author>
    <author><name>Chenyang Wan</name></author>
    <author><name>Jiaqi Peng</name></author>
    <author><name>Xiqian Yu</name></author>
    <author><name>Yuqiang Yang</name></author>
    <category term="Robotic Safety and Efficiency Systems"/>
    <summary type="text">DualVLN은 Vision-Language Navigation을 위해 고수준 추론(System 2)과 저수준 제어(System 1)를 분리한 최초의 dual-system foundation model으로, VLM 기반 global planner와 Diffusion Transformer 기반 policy의 비동기 협력을 통해 실시간 제어와 동적 장애물 회피를 가능하게 한다.</summary>
  </entry>
  <entry>
    <title>GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1411_GR-RL_Going_Dexterous_and_Precise_for_Long-Horizon_Robotic_M/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1411_GR-RL_Going_Dexterous_and_Precise_for_Long-Horizon_Robotic_M/</id>
    <updated>2025-12-01T00:00:00Z</updated>
    <author><name>Yunfei Li</name></author>
    <author><name>Xiao Ma</name></author>
    <author><name>Jiafeng Xu</name></author>
    <author><name>Yu Cui</name></author>
    <author><name>Zhongren Cui</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">GR-RL은 일반적인 vision-language-action (VLA) 정책을 다단계 학습 파이프라인(데이터 필터링, 형태 대칭 증강, 온라인 RL)을 통해 장기 복잡 조작을 위한 고정밀 전문가 정책으로 변환하는 로봇 학습 프레임워크이다.</summary>
  </entry>
  <entry>
    <title>An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1307_An_Anatomy_of_Vision-Language-Action_Models_From_Modules_to/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1307_An_Anatomy_of_Vision-Language-Action_Models_From_Modules_to/</id>
    <updated>2025-12-01T00:00:00Z</updated>
    <author><name>Chao Xu</name></author>
    <author><name>Suyu Zhang</name></author>
    <author><name>Yang Liu</name></author>
    <author><name>Baigui Sun</name></author>
    <author><name>Weihong Chen</name></author>
    <category term="Robotic Safety and Efficiency Systems"/>
    <summary type="text">Vision-Language-Action (VLA) 모델의 구조와 발전을 체계적으로 분석하는 종합 서베이로, 기본 모듈부터 역사적 마일스톤을 거쳐 5가지 핵심 과제까지 단계적으로 설명한다.</summary>
  </entry>
  <entry>
    <title>OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1500_OmniVLA_Physically-Grounded_Multimodal_VLA_with_Unified_Mult/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1500_OmniVLA_Physically-Grounded_Multimodal_VLA_with_Unified_Mult/</id>
    <updated>2025-11-01T00:00:00Z</updated>
    <author><name>Heyu Guo</name></author>
    <author><name>Shanmu Wang</name></author>
    <author><name>Ruichun Ma</name></author>
    <author><name>Shiqi Jiang</name></author>
    <author><name>Yasaman Ghasempour</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">OmniVLA는 RGB, 적외선, mmWave 레이더, 음향 마이크로폰 등 다중 센서를 통합하는 최초의 VLA 모델로, 센서-마스크된 이미지라는 통일된 표현을 통해 물리적 정보가 포함된 로봇 조작을 가능하게 한다.</summary>
  </entry>
  <entry>
    <title>NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1494_NORA-15_A_Vision-Language-Action_Model_Trained_using_World_M/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1494_NORA-15_A_Vision-Language-Action_Model_Trained_using_World_M/</id>
    <updated>2025-11-01T00:00:00Z</updated>
    <author><name>Chia-Yu Hung</name></author>
    <author><name>Navonil Majumder</name></author>
    <author><name>Haoyuan Deng</name></author>
    <author><name>Liu Renhang</name></author>
    <author><name>Yankang Ang</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">NORA-1.5는 flow-matching 기반 action expert를 추가하여 VLA 모델의 성능을 향상시키고, world model 및 action-based reward를 이용한 DPO 기반 post-training으로 실제 로봇 환경에서의 신뢰성과 일반화 능력을 개선한다.</summary>
  </entry>
  <entry>
    <title>IPR-1: Interactive Physical Reasoner</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1439_IPR-1_Interactive_Physical_Reasoner/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1439_IPR-1_Interactive_Physical_Reasoner/</id>
    <updated>2025-11-01T00:00:00Z</updated>
    <author><name>Mingyu Zhang</name></author>
    <author><name>Lifeng Zhuo</name></author>
    <author><name>Tianxi Tan</name></author>
    <author><name>Guocan Xie</name></author>
    <author><name>Xian Nie</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">Interactive Physical Reasoner (IPR)는 VLM의 정책을 world model의 롤아웃으로 강화하여 상호작용을 통해 물리 추론 능력을 학습하는 에이전트이다. PhysCode라는 물리 중심 액션 코드를 도입하여 의미론적 의도와 역학을 정렬하고, 1,000+ 게임으로 사전학습되어 물리 직관부터 목표 지향 추론까지 견고한 성능을 보인다.</summary>
  </entry>
  <entry>
    <title>GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion Policies</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1401_GauDP_Reinventing_Multi-Agent_Collaboration_through_Gaussian/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1401_GauDP_Reinventing_Multi-Agent_Collaboration_through_Gaussian/</id>
    <updated>2025-11-01T00:00:00Z</updated>
    <author><name>Ziye Wang</name></author>
    <author><name>Li Kang</name></author>
    <author><name>Yiran Qin</name></author>
    <author><name>Jiahua Ma</name></author>
    <author><name>Zhanglin Peng</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">GauDP는 다중 에이전트 협업 로봇 시스템에서 RGB 이미지로부터 3D Gaussian 필드를 구성하여 전역 일관성과 국소적 정밀성을 동시에 확보하는 새로운 표현 방식을 제안한다. 각 에이전트가 공유된 3D Gaussian 표현에서 과제 관련 특성을 동적으로 쿼리하여 협조와 개별 제어를 동시에 달성한다.</summary>
  </entry>
  <entry>
    <title>DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1373_DualVLA_Building_a_Generalizable_Embodied_Agent_via_Partial/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1373_DualVLA_Building_a_Generalizable_Embodied_Agent_via_Partial/</id>
    <updated>2025-11-01T00:00:00Z</updated>
    <author><name>Zhen Fang</name></author>
    <author><name>Zhuoyang Liu</name></author>
    <author><name>Jiaming Liu</name></author>
    <author><name>Hao Chen</name></author>
    <author><name>Yu Zeng</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">DualVLA는 Vision-Language-Action 모델에서 추론 능력을 추가할 때 발생하는 행동 성능 저하(action degeneration)를 해결하기 위해, 이중층 데이터 프루닝과 이중 교사 적응형 증류 전략을 통해 추론과 행동을 부분적으로 분리하는 접근법을 제시한다.</summary>
  </entry>
  <entry>
    <title>X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1633_X-VLA_Soft-Prompted_Transformer_as_Scalable_Cross-Embodiment/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1633_X-VLA_Soft-Prompted_Transformer_as_Scalable_Cross-Embodiment/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Jinliang Zheng</name></author>
    <author><name>Jianxiong Li</name></author>
    <author><name>Zhihao Wang</name></author>
    <author><name>Dongxiu Liu</name></author>
    <author><name>Xirui Kang</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">X-VLA는 소프트 프롬프트(Soft Prompt) 기법을 도입하여 이질적인 로봇 플랫폼 간 cross-embodiment 학습을 효과적으로 처리하는 scalable Vision-Language-Action 모델이다. 0.9B 파라미터 규모로 6개 시뮬레이션 벤치마크와 3개 실로봇에서 SOTA 성능을 달성한다.</summary>
  </entry>
  <entry>
    <title>World Simulation with Video Foundation Models for Physical AI</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1632_World_Simulation_with_Video_Foundation_Models_for_Physical_A/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1632_World_Simulation_with_Video_Foundation_Models_for_Physical_A/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Arslan Ali</name></author>
    <author><name>Junjie Bai</name></author>
    <author><name>Maciej Bala</name></author>
    <author><name>Yogesh Balaji</name></author>
    <author><name>Aaron Blakeman</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">Cosmos-Predict2.5는 flow-based architecture 기반의 세계 시뮬레이션 기초 모델로, Text2World, Image2World, Video2World 생성을 단일 모델에 통합하여 로보틱스와 자율주행 시스템을 위한 합성 데이터 생성과 폐루프 시뮬레이션을 가능하게 한다.</summary>
  </entry>
  <entry>
    <title>VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1619_VLA-RFT_Vision-Language-Action_Reinforcement_Fine-tuning_wit/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1619_VLA-RFT_Vision-Language-Action_Reinforcement_Fine-tuning_wit/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Hengtao Li</name></author>
    <author><name>Pengxiang Ding</name></author>
    <author><name>Runze Suo</name></author>
    <author><name>Yihao Wang</name></author>
    <author><name>Zirui Ge</name></author>
    <category term="LLM-Augmented Embodied Agent Frameworks"/>
    <summary type="text">VLA-RFT는 데이터 기반 world model을 시뮬레이터로 활용하여 vision-language-action 모델을 reinforcement learning으로 효율적으로 fine-tuning하는 프레임워크이다. 검증된 reward를 기반으로 GRPO 최적화를 수행하여 400 단계 이하의 fine-tuning으로 strong supervised baseline을 초과하는 성능을 달성한다.</summary>
  </entry>
  <entry>
    <title>VLA-0: Building State-of-the-Art VLAs with Zero Modification</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1615_VLA-0_Building_State-of-the-Art_VLAs_with_Zero_Modification/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1615_VLA-0_Building_State-of-the-Art_VLAs_with_Zero_Modification/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Ankit Goyal</name></author>
    <author><name>Hugo Hadfield</name></author>
    <author><name>Xuning Yang</name></author>
    <author><name>Valts Blukis</name></author>
    <author><name>Fabio Ramos</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">VLA-0는 Vision-Language Model의 구조 변경 없이 액션을 직접 텍스트로 표현하여 로봇 조작을 위한 최첨단 Vision-Language-Action 모델을 구축한다. 이 단순한 설계가 기존의 복잡한 방법들보다 우수한 성능을 달성한다.</summary>
  </entry>
  <entry>
    <title>Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1609_Vision-Language-Action_Models_for_Robotics_A_Review_Towards/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1609_Vision-Language-Action_Models_for_Robotics_A_Review_Towards/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Kento Kawaharazuka</name></author>
    <author><name>Jihoon Oh</name></author>
    <author><name>Jun Yamada</name></author>
    <author><name>Ingmar Posner</name></author>
    <author><name>Yuke Zhu</name></author>
    <category term="Robotic Safety and Efficiency Systems"/>
    <summary type="text">Vision-Language-Action (VLA) 모델이 로봇이 다양한 작업을 수행하도록 하는 통합 학습 방식에 대한 포괄적 리뷰로, 소프트웨어와 하드웨어 통합을 포함한 실제 배포 가이드를 제시한다.</summary>
  </entry>
  <entry>
    <title>TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1593_TrackVLA_Unleashing_Reasoning_and_Memory_Capabilities_in_VLA/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1593_TrackVLA_Unleashing_Reasoning_and_Memory_Capabilities_in_VLA/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Jiahang Liu</name></author>
    <author><name>Yunpeng Qi</name></author>
    <author><name>Jiazhao Zhang</name></author>
    <author><name>Minghan Li</name></author>
    <author><name>Shaoan Wang</name></author>
    <category term="Vision-Language Grounded Robot Navigation"/>
    <summary type="text">TrackVLA++는 Vision-Language-Action 모델에 Polar-CoT 공간 추론과 Target Identification Memory(TIM)를 통합하여 장시간 추적과 폐색 상황에서의 강건한 embodied visual tracking을 실현한다.</summary>
  </entry>
  <entry>
    <title>Running VLAs at Real-time Speed</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1557_Running_VLAs_at_Real-time_Speed/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1557_Running_VLAs_at_Real-time_Speed/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Yunchao Ma</name></author>
    <author><name>Yizhuang Zhou</name></author>
    <author><name>Yunhuan Yang</name></author>
    <author><name>Tiancai Wang</name></author>
    <author><name>Haoqiang Fan</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">π0 레벨의 multi-view VLA를 단일 소비자 GPU에서 30Hz 프레임 레이트로 실행하기 위해 모델 추론 오버헤드를 제거하는 최적화 기법들을 제시하고, 실시간 로봇 제어를 위한 Full Streaming Inference 프레임워크를 제안한다.</summary>
  </entry>
  <entry>
    <title>RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1532_RLinf-VLA_A_Unified_and_Efficient_Framework_for_Reinforcemen/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1532_RLinf-VLA_A_Unified_and_Efficient_Framework_for_Reinforcemen/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Hongzhi Zang</name></author>
    <author><name>Mingjie Wei</name></author>
    <author><name>Si Xu</name></author>
    <author><name>Yongji Wu</name></author>
    <author><name>Zhen Guo</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">RLinf-VLA는 Vision-Language-Action 모델의 강화학습 훈련을 위한 통합되고 효율적인 프레임워크로, 다양한 VLA 아키텍처, RL 알고리즘, 시뮬레이터를 지원하며 GPU 할당 최적화를 통해 2.27배 속도 향상을 달성한다.</summary>
  </entry>
  <entry>
    <title>InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1438_InternVLA-M1_A_Spatially_Guided_Vision-Language-Action_Frame/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1438_InternVLA-M1_A_Spatially_Guided_Vision-Language-Action_Frame/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Xinyi Chen</name></author>
    <author><name>Yilun Chen</name></author>
    <author><name>Yanwei Fu</name></author>
    <author><name>Ning Gao</name></author>
    <author><name>Jiaya Jia</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">InternVLA-M1은 공간 그라운딩을 시각-언어-행동 학습의 중심 연결고리로 활용하여, 지시 따르기 로봇의 확장 가능한 일반 지능을 구현한 통합 프레임워크이다.</summary>
  </entry>
  <entry>
    <title>Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1403_Gemini_Robotics_15_Pushing_the_Frontier_of_Generalist_Robots/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1403_Gemini_Robotics_15_Pushing_the_Frontier_of_Generalist_Robots/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Gemini Robotics Team</name></author>
    <author><name>Abbas Abdolmaleki</name></author>
    <author><name>Saminda Abeyruwan</name></author>
    <author><name>Joshua Ainslie</name></author>
    <author><name>Jean-Baptiste Alayrac</name></author>
    <category term="3D Simulation and Robot Manipulation"/>
    <summary type="text">Gemini Robotics 1.5는 Motion Transfer 메커니즘과 embodied thinking 능력을 통해 다중 로봇 플랫폼을 제어할 수 있는 Vision-Language-Action 모델이며, Gemini Robotics-ER 1.5는 embodied reasoning에서 최첨단 성능을 달성하는 Vision-Language 모델이다.</summary>
  </entry>
  <entry>
    <title>D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1347_D2E_Scaling_Vision-Action_Pretraining_on_Desktop_Data_for_Tr/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1347_D2E_Scaling_Vision-Action_Pretraining_on_Desktop_Data_for_Tr/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Suhwan Choi</name></author>
    <author><name>Jaeyoon Jung</name></author>
    <author><name>Haebin Seong</name></author>
    <author><name>Minchan Kim</name></author>
    <author><name>Minyeong Kim</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">D2E는 데스크톱 환경(게임 등)에서 수집한 대규모 비전-액션 데이터를 사전학습 자료로 사용하여 로봇 조작 및 네비게이션 같은 구체화된 AI 작업으로 전이 학습하는 프레임워크를 제시한다.</summary>
  </entry>
  <entry>
    <title>Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1337_Compose_Your_Policies_Improving_Diffusion-based_or_Flow-base/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1337_Compose_Your_Policies_Improving_Diffusion-based_or_Flow-base/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Jiahang Cao</name></author>
    <author><name>Yize Huang</name></author>
    <author><name>Hanzhong Guo</name></author>
    <author><name>Rui Zhang</name></author>
    <author><name>Mu Nan</name></author>
    <category term="LLM-Augmented Embodied Agent Frameworks"/>
    <summary type="text">본 논문은 General Policy Composition (GPC)를 제안하여 사전학습된 diffusion 또는 flow 기반 로봇 정책들의 분포 수준 점수를 convex 조합으로 결합함으로써, 추가 학습 없이 개별 정책보다 우수한 성능을 달성한다.</summary>
  </entry>
  <entry>
    <title>A Comprehensive Survey on World Models for Embodied AI</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1292_A_Comprehensive_Survey_on_World_Models_for_Embodied_AI/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1292_A_Comprehensive_Survey_on_World_Models_for_Embodied_AI/</id>
    <updated>2025-10-01T00:00:00Z</updated>
    <author><name>Xinqing Li</name></author>
    <author><name>Xin He</name></author>
    <author><name>Le Zhang</name></author>
    <author><name>Min Wu</name></author>
    <author><name>Xiaoli Li</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">Embodied AI를 위한 World Models에 대한 포괄적 조사로, Functionality, Temporal Modeling, Spatial Representation의 세 축 분류체계를 제안하여 환경 동역학을 캡처하고 예측하는 내부 시뮬레이터를 체계적으로 정리한다.</summary>
  </entry>
  <entry>
    <title>VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1618_VLA-Reasoner_Empowering_Vision-Language-Action_Models_with_R/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1618_VLA-Reasoner_Empowering_Vision-Language-Action_Models_with_R/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Wenkai Guo</name></author>
    <author><name>Guanxing Lu</name></author>
    <author><name>Haoyuan Deng</name></author>
    <author><name>Zhenyu Wu</name></author>
    <author><name>Yansong Tang</name></author>
    <category term="Robotic Safety and Efficiency Systems"/>
    <summary type="text">VLA-Reasoner는 Vision-Language-Action 모델에 test-time MCTS를 통합하여 장기 지평 로봇 조작 작업에서 누적 편차를 해결하고 미래 상태를 예측하는 플러그인 프레임워크이다.</summary>
  </entry>
  <entry>
    <title>VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1616_VLA-Adapter_An_Effective_Paradigm_for_Tiny-Scale_Vision-Lang/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1616_VLA-Adapter_An_Effective_Paradigm_for_Tiny-Scale_Vision-Lang/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Yihao Wang</name></author>
    <author><name>Pengxiang Ding</name></author>
    <author><name>Lingxiao Li</name></author>
    <author><name>Can Cui</name></author>
    <author><name>Zirui Ge</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">VLA-Adapter는 경량 백본(0.5B 파라미터)을 사용하여 로봇 데이터 사전학습 없이 최첨단 Vision-Language-Action 모델을 학습할 수 있는 새로운 패러다임을 제시한다. Bridge Attention을 통해 비전-언어 표현을 행동 공간에 효과적으로 연결한다.</summary>
  </entry>
  <entry>
    <title>SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1577_SpecPrune-VLA_Accelerating_Vision-Language-Action_Models_via/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1577_SpecPrune-VLA_Accelerating_Vision-Language-Action_Models_via/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Hanzhen Wang</name></author>
    <author><name>Jiaming Xu</name></author>
    <author><name>Yushun Xiang</name></author>
    <author><name>Jiayi Pan</name></author>
    <author><name>Yongkang Zhou</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">SpecPrune-VLA는 Vision-Language-Action 모델의 LLM 추론을 가속화하기 위해 시간-공간 일관성을 활용한 액션-인식 자체-추측 토큰 프루닝 기법을 제안한다. 두 단계 프루닝(액션 레벨 정적 프루닝과 레이어 레벨 동적 프루닝)과 액션-인식 컨트롤러를 통해 최대 1.70배 속도 향상을 달성한다.</summary>
  </entry>
  <entry>
    <title>SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1573_SimpleVLA-RL_Scaling_VLA_Training_via_Reinforcement_Learning/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1573_SimpleVLA-RL_Scaling_VLA_Training_via_Reinforcement_Learning/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Haozhan Li</name></author>
    <author><name>Yuxin Zuo</name></author>
    <author><name>Jiale Yu</name></author>
    <author><name>Yuhao Zhang</name></author>
    <author><name>Zhaohui Yang</name></author>
    <category term="3D Simulation and Robot Manipulation"/>
    <summary type="text">SimpleVLA-RL은 Vision-Language-Action 모델의 학습을 강화학습(RL)을 통해 확장하는 효율적인 프레임워크로, 데이터 부족 문제를 해결하고 실제 로봇 작업에서 SFT를 능가하는 성능을 달성한다.</summary>
  </entry>
  <entry>
    <title>Pure Vision Language Action (VLA) Models: A Comprehensive Survey</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1519_Pure_Vision_Language_Action_VLA_Models_A_Comprehensive_Surve/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1519_Pure_Vision_Language_Action_VLA_Models_A_Comprehensive_Surve/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Dapeng Zhang</name></author>
    <author><name>Jing Sun</name></author>
    <author><name>Chenghui Hu</name></author>
    <author><name>Xiaoyan Wu</name></author>
    <author><name>Zhenlong Yuan</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">본 논문은 Vision Language Action (VLA) 모델을 체계적으로 분류하고 분석하는 포괄적 서베이로, autoregression-based, diffusion-based, reinforcement-based, hybrid, specialized methods로 VLA 접근법을 분류하여 300개 이상의 최근 연구를 종합한다.</summary>
  </entry>
  <entry>
    <title>OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1499_OmniVLA_An_Omni-Modal_Vision-Language-Action_Model_for_Robot/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1499_OmniVLA_An_Omni-Modal_Vision-Language-Action_Model_for_Robot/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Noriaki Hirose</name></author>
    <author><name>Catherine Glossop</name></author>
    <author><name>Dhruv Shah</name></author>
    <author><name>Sergey Levine</name></author>
    <category term="Vision-Language Grounded Robot Navigation"/>
    <summary type="text">OmniVLA는 2D 포즈, egocentric 이미지, 자연어 등 다양한 모달리티로 조건화된 목표를 처리할 수 있는 omni-modal vision-language-action 모델로, 9,500시간 이상의 다중 플랫폼 로봇 네비게이션 데이터로 학습되어 강력한 일반화 성능을 달성한다.</summary>
  </entry>
  <entry>
    <title>ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1465_ManiFlow_A_General_Robot_Manipulation_Policy_via_Consistency/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1465_ManiFlow_A_General_Robot_Manipulation_Policy_via_Consistency/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Ge Yan</name></author>
    <author><name>Jiyue Zhu</name></author>
    <author><name>Yuquan Deng</name></author>
    <author><name>Shiqi Yang</name></author>
    <author><name>Ri-Zhao Qiu</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">ManiFlow는 flow matching과 consistency training을 결합하여 1-2 inference step으로 고품질의 dexterous action을 생성하는 visuomotor imitation learning policy이다. DiT-X 아키텍처를 통해 visual, language, proprioceptive 입력을 효율적으로 조건화하며 실제 로봇 환경에서 우수한 성능을 보인다.</summary>
  </entry>
  <entry>
    <title>JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1441_JanusVLN_Decoupling_Semantics_and_Spatiality_with_Dual_Impli/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1441_JanusVLN_Decoupling_Semantics_and_Spatiality_with_Dual_Impli/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Shuang Zeng</name></author>
    <author><name>Dekang Qi</name></author>
    <author><name>Xinyuan Chang</name></author>
    <author><name>Feng Xiong</name></author>
    <author><name>Shichao Xie</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">JanusVLN은 시각-언어 네비게이션에서 spatial-geometric과 visual-semantic 정보를 분리하여 dual implicit neural memory로 모델링하는 프레임워크를 제안한다. 3D 기하학적 선행 지식과 MLLM의 의미론적 이해를 결합하여 효율적이고 공간 인식적인 에이전트 네비게이션을 실현한다.</summary>
  </entry>
  <entry>
    <title>GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1402_GC-VLN_Instruction_as_Graph_Constraints_for_Training-free_Vi/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1402_GC-VLN_Instruction_as_Graph_Constraints_for_Training-free_Vi/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Hang Yin</name></author>
    <author><name>Haoyu Wei</name></author>
    <author><name>Xiuwei Xu</name></author>
    <author><name>Wenxuan Guo</name></author>
    <author><name>Jie Zhou</name></author>
    <category term="Vision-Language Grounded Robot Navigation"/>
    <summary type="text">GC-VLN은 자연언어 지시를 그래프 제약 최적화 문제로 재구성하여 연속 환경에서 학습 없이 작동하는 비전-언어 네비게이션 프레임워크를 제안한다. 공간 제약 라이브러리와 제약 솔버를 통해 zero-shot 환경 적응을 실현한다.</summary>
  </entry>
  <entry>
    <title>Embodied Navigation Foundation Model</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1378_Embodied_Navigation_Foundation_Model/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1378_Embodied_Navigation_Foundation_Model/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Jiazhao Zhang</name></author>
    <author><name>Anqi Li</name></author>
    <author><name>Yunpeng Qi</name></author>
    <author><name>Minghan Li</name></author>
    <author><name>Jiahang Liu</name></author>
    <category term="Vision-Language Grounded Robot Navigation"/>
    <summary type="text">NavFoM은 8백만 개의 네비게이션 샘플로 학습된 크로스-구현체·크로스-태스크 기반 네비게이션 모델로, 다양한 로봇 플랫폼과 네비게이션 작업에서 미세 조정 없이 최첨단 성능을 달성한다.</summary>
  </entry>
  <entry>
    <title>Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1346_Cross-Platform_Scaling_of_Vision-Language-Action_Models_from/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1346_Cross-Platform_Scaling_of_Vision-Language-Action_Models_from/</id>
    <updated>2025-09-01T00:00:00Z</updated>
    <author><name>Amir Taherin</name></author>
    <author><name>Juyi Lin</name></author>
    <author><name>Arash Akbari</name></author>
    <author><name>Arman Akbari</name></author>
    <author><name>Pu Zhao</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">Vision-Language-Action (VLA) 모델의 성능을 엣지 디바이스부터 데이터센터 GPU까지 다양한 하드웨어 플랫폼에서 체계적으로 평가하여, 아키텍처와 하드웨어 제약 조건에 따른 정확도, 레이턴시, 처리량, 메모리 사용량의 확장 추이를 밝혀낸다.</summary>
  </entry>
  <entry>
    <title>Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1446_Large_VLM-based_Vision-Language-Action_Models_for_Robotic_Ma/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1446_Large_VLM-based_Vision-Language-Action_Models_for_Robotic_Ma/</id>
    <updated>2025-08-01T00:00:00Z</updated>
    <author><name>Rui Shao</name></author>
    <author><name>Wei Li</name></author>
    <author><name>Lingsen Zhang</name></author>
    <author><name>Renshan Zhang</name></author>
    <author><name>Zhiyang Liu</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">대규모 Vision-Language Model(VLM)을 기반으로 한 Vision-Language-Action(VLA) 모델들을 로봇 매니퓰레이션에 적용하는 연구의 첫 번째 체계적 설문조사로, Monolithic 모델과 Hierarchical 모델이라는 두 가지 주요 아키텍처 패러다임을 제시한다.</summary>
  </entry>
  <entry>
    <title>Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1445_Large_Model_Empowered_Embodied_AI_A_Survey_on_Decision-Makin/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1445_Large_Model_Empowered_Embodied_AI_A_Survey_on_Decision-Makin/</id>
    <updated>2025-08-01T00:00:00Z</updated>
    <author><name>Wenlong Liang</name></author>
    <author><name>Rui Zhou</name></author>
    <author><name>Yang Ma</name></author>
    <author><name>Bing Zhang</name></author>
    <author><name>Songlin Li</name></author>
    <category term="LLM-Augmented Embodied Agent Frameworks"/>
    <summary type="text">대규모 모델이 강화된 embodied AI 시스템의 의사결정과 학습 방법을 체계적으로 조사한 종합 서베이로, 계층적/end-to-end 의사결정 패러다임, imitation learning/reinforcement learning 기반 embodied learning, 그리고 world model의 역할을 통합적으로 분석한다.</summary>
  </entry>
  <entry>
    <title>Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1406_Genie_Envisioner_A_Unified_World_Foundation_Platform_for_Rob/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1406_Genie_Envisioner_A_Unified_World_Foundation_Platform_for_Rob/</id>
    <updated>2025-08-01T00:00:00Z</updated>
    <author><name>Yue Liao</name></author>
    <author><name>Pengfei Zhou</name></author>
    <author><name>Siyuan Huang</name></author>
    <author><name>Donglin Yang</name></author>
    <author><name>Shengcong Chen</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">Genie Envisioner는 video diffusion model 기반의 통합 로봇 조작 플랫폼으로, 정책 학습, 평가, 시뮬레이션을 단일 비디오 생성 프레임워크 내에서 통합한다.</summary>
  </entry>
  <entry>
    <title>EO-1: An Open Unified Embodied Foundation Model for General Robot Control</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1385_EO-1_An_Open_Unified_Embodied_Foundation_Model_for_General_R/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1385_EO-1_An_Open_Unified_Embodied_Foundation_Model_for_General_R/</id>
    <updated>2025-08-01T00:00:00Z</updated>
    <author><name>Delin Qu</name></author>
    <author><name>Haoming Song</name></author>
    <author><name>Qizhi Chen</name></author>
    <author><name>Zhaoqing Chen</name></author>
    <author><name>Xianqiang Gao</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">EO-1은 interleaved vision-text-action 사전학습을 통해 multimodal embodied reasoning과 robot control을 통합한 unified embodied foundation model이며, 1.5M 샘플의 EO-Data1.5M 데이터셋과 함께 개발되었다.</summary>
  </entry>
  <entry>
    <title>Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1380_Embodied-R1_Reinforced_Embodied_Reasoning_for_General_Roboti/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1380_Embodied-R1_Reinforced_Embodied_Reasoning_for_General_Roboti/</id>
    <updated>2025-08-01T00:00:00Z</updated>
    <author><name>Yifu Yuan</name></author>
    <author><name>Haiqin Cui</name></author>
    <author><name>Yaoting Huang</name></author>
    <author><name>Yibin Chen</name></author>
    <author><name>Fei Ni</name></author>
    <category term="LLM-Augmented Embodied Agent Frameworks"/>
    <summary type="text">Embodied-R1은 '포인팅'을 통일된 embodiment-agnostic 중간 표현으로 정의하고, Reinforced Fine-tuning(RFT)으로 훈련된 3B VLM으로서 로봇 조작의 perception-action gap을 효과적으로 극복한다.</summary>
  </entry>
  <entry>
    <title>DiWA: Diffusion Policy Adaptation with World Models</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1368_DiWA_Diffusion_Policy_Adaptation_with_World_Models/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1368_DiWA_Diffusion_Policy_Adaptation_with_World_Models/</id>
    <updated>2025-08-01T00:00:00Z</updated>
    <author><name>Akshay L Chandra</name></author>
    <author><name>Iman Nematollahi</name></author>
    <author><name>Chenguang Huang</name></author>
    <author><name>Tim Welschehold</name></author>
    <author><name>Wolfram Burgard</name></author>
    <category term="VLA Policy Training and Adaptation"/>
    <summary type="text">DiWA는 학습된 world model을 활용하여 diffusion 기반 로봇 정책을 오프라인으로 미세조정하는 프레임워크로, RL을 통해 상상 속 롤아웃에서 정책을 개선한다.</summary>
  </entry>
  <entry>
    <title>Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies</title>
    <link href="https://paper-curation.jehyunlee.dev/papers/1366_Discrete_Diffusion_VLA_Bringing_Discrete_Diffusion_to_Action/" rel="alternate" type="text/html"/>
    <id>https://paper-curation.jehyunlee.dev/papers/1366_Discrete_Diffusion_VLA_Bringing_Discrete_Diffusion_to_Action/</id>
    <updated>2025-08-01T00:00:00Z</updated>
    <author><name>Zhixuan Liang</name></author>
    <author><name>Yizhuo Li</name></author>
    <author><name>Tianshuo Yang</name></author>
    <author><name>Chengyue Wu</name></author>
    <author><name>Sitong Mao</name></author>
    <category term="Vision-Language-Action Model Architectures"/>
    <summary type="text">Vision-Language-Action (VLA) 모델에 discrete diffusion을 적용하여 action token을 적응적으로 디코딩하는 unified transformer 정책을 제시한다. 이를 통해 자동회귀 방식의 순서 제약을 극복하고 분리된 decoder 구조의 문제를 해결한다.</summary>
  </entry>
</feed>
