저자: Semih Kara, Oguzhan Ersoy | 날짜: 2026 | URL: https://openreview.net/forum?id=KNwvVyPs2N 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Figure 3. Fully-correct rollout. Per-token advantage ASD
reward model 학습 없이 self-distillation의 context로 step-aligned critique를 사용하면, PRM처럼 오류가 발생한 step에 정확히 국소화된 per-token advantage를 만들어낼 수 있음을 보인 연구이다.
Figure 2. STEPALIGNFB matches or exceeds GRPO and REFSOL on OpenMathReasoning. REFSOL conditions the teacher on the
Figure 4. Incorrect step. Per-token advantages for a rollout containing an arithmetic/reasoning error at the step marked
총평: reward model 학습 없이도 feedback을 solver trace에 정렬시키는 것만으로 process supervision과 유사한 국소화된 credit assignment를 얻을 수 있음을 명확한 실험 설계와 advantage 분석으로 설득력 있게 보여준 워크숍급 논문으로, 실용적 파급력이 크지만 검증 범위가 제한적이다.