RetrOrchestrator: A Multi-Step Retrosynthesis Agent Dynamically Orchestrating Single-Step Transition Models

저자: Liao Chang, Luotian Yuan, Yiping Ke, Ying Wei | 날짜: 2026 | URL: https://openreview.net/forum?id=p6gN6f8pdy 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 2

Figure 2. Comparison of our proposed methodology, formulated

RetrOrchestrator는 multi-step retrosynthesis planning을 POMDP로 재정의하여, 각 단계마다 확장할 분자뿐 아니라 어떤 single-step retrosynthesis model(SSR)을 tool로 사용할지 동적으로 선택하는 LLM 기반 agent이다. scaffold-aware 강화학습(SA-GRPO)을 통해 sparse reward 문제를 해결하며 Retro*-190과 PDB-600에서 state-of-the-art 성공률을 달성한다.

Motivation

Achievement

Figure 1

Figure 1. Relative performance of retrosynthesis agents with dif-

  1. 최초의 LLM 기반 retrosynthesis planning agent: multi-step retrosynthesis를 POMDP로 정식화하여 LLM의 context를 implicit belief state로 활용, 여러 SSR model을 동적으로 orchestrate하는 최초의 agent를 제안했다.
  2. SA-GRPO 알고리즘: scaffold-level clustering을 이용해 sparse하고 delayed된 reward 상황에서도 각 SSR model에 대한 credit assignment를 가능하게 하고, 쉬운 분자에서의 chemical rule 학습부터 어려운 분자에서의 전략적 planning까지 점진적으로 전이하는 curriculum을 설계했다.
  3. Pareto-optimal 성능 달성: Retro*-190 벤치마크에서 94.21%의 state-of-the-art 성공률(off-the-shelf LLM 9.47%, non-LLM state-dependent router 82.63% 대비)을 달성했으며, 해결된 route의 92.49%가 2개 이상의 SSR을 호출해 정책이 단일 specialist나 static router로 붕괴하지 않았음을 입증했다. PDB-600이라는 더 큰 out-of-distribution set에서도 동일한 성능 향상이 유지되며, wall-clock time과 model-query count 양 측면에서 Pareto-optimal을 달성했다.

How

Figure 3

Figure 3. Overview of the RetrOrchestrator training algorithm and pipeline. The bottom panel illustrates the scaffold-aw

Originality

Limitation & Further Study

Evaluation

Novelty: 5/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: SSR model 간 skill disparity라는 실질적으로 중요하지만 간과되어온 문제를 POMDP 정식화와 scaffold-aware 강화학습으로 해결한 독창적이고 실용적인 연구로, retrosynthesis planning 분야에 새로운 방향을 제시한다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'PRIME: A Multi-Agent Environment for Orchestrating Dynamic Computational Workflows in Protein Engineerings'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구POMDP 기반 계획의 이론적 토대를 제공한다.
기반 연구구조화된 중간 상태 기반 화학 반응 예측을 확장한다.
다른 접근multi-step retrosynthesis planning을 위한 다른 방법론을 제시한다.
기반 연구STM 기반 조립 계획 프레임워크를 확장한다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근single-step retrosynthesis model 선택을 위한 다른 접근을 다룬다.
다른 접근single-step retrosynthesis model 선택 문제와 관련된 대안 접근이다.
다른 접근retrosynthesis 예측 모델을 도구로 활용하는 유사 프레임워크이다.
다른 접근POMDP 기반 순차적 의사결정을 다루는 유사 연구이다.
후속 연구동적 tool 선택 방식을 확장한 연구이다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드