저자: Yuhan Chi | 날짜: 2026 | URL: https://openreview.net/forum?id=wRImV3kfR1 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Figure 1: Overview of core findings. (a) Median AUROC under three evaluation protocols. (b)
본 논문은 token-level entropy/log-probability 기반 LLM reasoning 검증기의 AUROC가 널리 보고되는 0.55~0.80+라는 편차가 실제 신호의 차이보다는 evaluation protocol의 차이(특히 global pooling, in-sample leakage, direction-agnostic scoring)에서 비롯됨을 controlled experiment로 규명한다.
Figure 1: Overview of core findings. (a) Median AUROC under three evaluation protocols. (b)
Figure 1: Overview of core findings. (a) Median AUROC under three evaluation protocols. (b)
총평: 새로운 알고리즘을 제안하기보다 evaluation protocol의 함정을 체계적으로 드러내는 메타적 기여를 하는 논문으로, token-level verification 연구 커뮤니티에 실질적인 재현성 및 평가 표준 개선에 기여할 수 있는 실용적 가치가 크다.