저자: Dulhan Jayalath, Shashwat Goel, Thomas Foster, Parag Jain, Suchin Gururangan, Cheng Zhang, Anirudh Goyal, Alan Schelten | 날짜: 2026 | URL: https://openreview.net/forum?id=oWOYBP7uFd 📄 PDF
Essence
Figure 1. CaT framework. The policy πt generates G rollouts for
추론 시점의 병렬 rollout 계산을 정답(reference)이 없는 post-training 환경에서 학습 신호로 변환하는 프레임워크인 Compute as Teacher (CaT)를 제안하며, reference estimation(synthesis)과 reward derivation(self-proposed rubrics)이라는 두 단계를 통해 검증 불가능한 도메인에서도 RL 학습을 가능케 한다.
Evaluation
Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5
총평: 인간 라벨이나 도메인 verifier 없이도 검증 불가능한 실세계 도메인에서 RL 학습을 가능케 하는 실용적이고 확장 가능한 프레임워크로, self-proposed rubrics라는 아이디어가 특히 신선하며 의료 도메인에서의 강력한 실증 결과가 인상적이다. 다만 LLM judge 의존성과 rollout 수렴에 따른 신호 약화 문제는 추가 검증이 필요하다.
같이 보면 좋은 논문
기반 연구SPECTER2 유사도 0.92로 LLM Agent Reasoning Training와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Evaluation of openai o1: Opportunities and challenges of agi'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 LLM Agent Reasoning Training와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 LLM Agent Reasoning Training와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구inference compute를 학습 신호로 전환하는 이론적 기초를 제공함
다른 접근medical VLM의 RL 보상 설계에 대한 다른 접근법을 제시하는 연구로 보인다.
기반 연구분포 변화 하의 안전성 검증을 확장한다.
기반 연구conformal prediction을 다른 응용 문제에 적용한 연구이다.
응용 사례추론 계산 활용을 수학적 유도 검증 문제에 적용한 사례이다.
기반 연구privileged information 기반 학습을 실제 배포 모델에 적용한 사례이다.
다른 접근distribution-aware 최적 정책 설계라는 이론적 기반을 공유한다.
다른 접근post-training에서 학습 신호를 생성하는 다른 접근법
후속 연구e-value/martingale 기반 anytime-valid 검정의 이론적 기반을 제공하는 연구이다.
다른 접근reference-free post-training 신호 생성을 위한 대안적 방법을 제시한다.
후속 연구risk control 이론적 기반을 제공함
후속 연구policy gradient 기반 reasoning 학습의 이론적 토대를 제공한다.
후속 연구anytime-valid 신뢰구간의 이론적 기반을 제공한다.
응용 사례reference-free 환경에서 reward 신호 생성을 실제 post-training에 적용한 연구
후속 연구RLVR의 보상 게이팅 메커니즘에 대한 기초 연구
반론/비판confidence 기반 pseudo-labeling의 한계에 대해 반박하는 관점을 제시한다.