Labimus: A Simulation and Benchmark for Humanoid Dexterous Manipulation in Chemical Laboratory
저자: Yuhan Wu, Zhao Jin, Tao Li, Yuheng Zhang, Zhichao Wang, Shuo Wang, Jun Jiang, Xiaobo Li, Yanyong Zhang, Jian Tang, Zhengping Che, Yan Xia | 날짜: 2026-07-01 | DOI: 10.48550/arXiv.2606.31037
⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.
Essence
Figure 1: Overview of Labimus. Top: real-to-sim reconstruction of a chemistry workstation with
Labimus는 화학 실험실(특히 유기화학) 환경에서 humanoid dexterous manipulation을 평가하는 최초의 benchmark로, real-to-sim 방식으로 재구성한 30여 개의 정밀 실험 기자재 asset과 particle-based powder physics, closed-loop instrument readout을 통합하여 조작-측정 파이프라인 전체를 검증한다.
Motivation
Known: 기존 로봇 실험실 자동화 연구는 fixed robotic arm과 task-specific gripper를 이용해 정형화된 환경에서 predefined experiment를 수행하는 방식에 집중해왔으며, RLBench, CALVIN, RoboCasa, LIBERO, BEHAVIOR-1K, Genie Sim 3.0, ManiSkill3 등 humanoid 및 dexterous manipulation benchmark들은 household task 중심으로 binary success 기준을 사용한다.
Gap: solid-solid transfer와 같은 정밀 조작은 재료와 조건에 따라 실시간 적응이 필요해 fixed-arm 시스템으로 표준화하기 어렵고, humanoid dexterous hand를 활용할 필요가 있음에도 이를 precision-critical laboratory 환경에서 평가하는 benchmark가 전무했다.
Why: 실험실 작업은 milligram 단위 정밀도, multi-finger coordination, 지속적 tool-mediated contact를 요구하는데, 기존 binary success 기준으로는 task completion과 experimental validity 간의 괴리를 드러낼 수 없으므로, 과학 실험 자동화를 위한 신뢰할 수 있는 humanoid robot 개발에 이러한 정밀 평가 체계가 필수적이다.
Approach: Labimus는 실제 유기화학 workstation을 real-to-sim reconstruction하여 30개 이상의 기능적으로 충실한 asset을 만들고, 6개 atomic operation과 7단계 solid-weighing workflow를 정의하며, task completion·experimental precision·long-horizon execution을 함께 측정하는 precision-aware evaluation protocol을 제안한다.
Achievement
Figure 4: Solid-weighing task suite and evaluation conditions. (a) The task suite spans three
최초의 정밀 실험실 humanoid manipulation benchmark: 유기화학 실험실에서 6개 atomic operation과 실제 SOP 기반 7단계 solid-weighing workflow를 포함하는 benchmark를 처음으로 구축함.
고충실도 실험실 시뮬레이션: articulated laboratory instrument, particle-based powder physics, closed-loop instrument readout(예: 실시간 디지털 mass display)을 통합해 manipulation-to-measurement pipeline을 완성함.
정밀도 인지 평가 프로토콜: binary task completion, tolerance 기반 quantitative precision(예: ±0.001 g), step-level/stage-level long-horizon execution을 결합한 3-tier 평가 체계와 조명·질감 perturbation을 포함한 3×4 evaluation matrix를 도입함.
정밀도 격차(precision gap) 발견: ACT, Diffusion Policy, π0 세 가지 대표 정책을 벤치마킹한 결과, task를 성공적으로 완료해도 실험 프로토콜이 요구하는 정량적 tolerance를 만족하지 못하는 경우가 존재함을 실증함.
How
Figure 3: Labimus benchmark overview. Top-left: the simulation environment in Isaac Sim with
실제 유기화학 workstation을 촬영/스캔하여 flask, beaker, spatula, analytical balance 등 30개 이상의 기능적 asset을 real-to-sim 방식으로 재구성
각 asset에 기능에 따라 articulated mechanism, particle-based powder physics, closed-loop instrument readout 중 해당하는 요소를 부여
실제 SOP를 기반으로 door open, door close, grasp & place, tare press, tool pickup, scoop & weigh의 6개 atomic operation과 7단계 solid-weighing workflow 정의
Isaac Gym 기반 시뮬레이션 환경에서 procedural layout과 lighting/texture perturbation(및 결합) 조건 하에 3×4 evaluation matrix 구성
ACT, Diffusion Policy, π0 세 가지 policy를 학습/평가하여 binary completion, tolerance 기반 precision, step/stage 단위 long-horizon execution을 측정
Originality
기존 household 중심 humanoid/dexterous benchmark(RLBench, CALVIN, RoboCasa, LIBERO, BEHAVIOR-1K, Genie Sim 3.0, ManiSkill3 등)와 달리 chemistry/biology 같은 precision-critical laboratory 도메인을 다룬 최초 사례
particle-based powder physics와 closed-loop instrument readout을 결합하여 manipulation 결과를 실제 정량적 측정치로 검증 가능하게 한 real-to-sim 파이프라인
binary success가 아닌 quantitative tolerance 기반 precision metric과 long-horizon step/stage diagnostic을 결합한 다층적 평가 프로토콜 설계
Limitation & Further Study
본문 발췌만으로는 particle-based powder physics의 물리적 정확도(실제 powder와의 sim-to-real gap) 검증이 충분히 제시되지 않아 추가 검증이 필요
벤치마킹된 policy가 ACT, Diffusion Policy, π0 세 가지로 제한적이며, 더 다양한 최신 humanoid manipulation policy 및 실제 로봇(Tianyi 2.0) 상에서의 sim-to-real 전이 검증이 후속 연구로 필요
유기화학 solid-weighing workflow 중심으로 설계되어, 액체 취급(liquid handling)이나 다른 실험실 도메인(생물학 등)으로의 일반화 가능성은 추가 확장이 필요
총평: Labimus는 humanoid dexterous manipulation 연구를 household 중심에서 정밀도가 중요한 과학 실험실 도메인으로 확장한 의미 있는 최초의 benchmark로, task completion과 experimental validity 간의 근본적 괴리를 명확히 드러낸 점에서 학술적·실용적 가치가 크다.