저자: Brando Miranda, Srivatsava Daruru, Ethan S Hersch, Zhanke Zhou, Allen Nie, Daneshvar Amrollahi, Leni Aniva, Iddah Mlauzi, Kirill Acharya, Elyas Obbad, Dilara Soylu, Weston Kirk, Zixiao Jolene Wang, Kai Fronsdal, Ying Li, Donald Poindexter Jr, Rakshit Kaushik, Shurui Liu, Yegor Denisov-Blanch, Steven Dillmann, Simon Obstbaum, Santiago Cuellar, John Sarracino, Rylan Schaeffer, Mo Tiwari, Donghyun Lee, Bo Han, Sanmi Koyejo | 날짜: 2026 | URL: https://openreview.net/forum?id=lkL0qnUv3p 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Figure 2. VERIBENCH evaluation topology. Each solid arrow is
VeriBench는 Python 소스 코드에서 Lean 4 형식 검증 아티팩트로의 end-to-end autoformalization을 평가하는 896개 과제 벤치마크로, SCSC(Smooth Conjunctive Score for Code verification)라는 5개 요소의 로그 도메인 기하평균 지표를 통해 typecheck, sorry 없는 증명, reference theorem과의 의미적 커버리지, reference-side validity gate를 결합 평가한다.
Figure 3. Human score distributions (0–5): lenient set (n = 297, left) and expert/harsh set (n = 193, right).
Figure 2. VERIBENCH evaluation topology. Each solid arrow is
총평: 테스트 기반 평가의 구조적 한계를 넘어 end-to-end formal verification을 agentic 방식으로 평가하는 새로운 벤치마크와 지표를 제시하며, specification synthesis가 proof search 못지않은 병목임을 보인 의미 있는 실증 연구이나 coverage 판정이 LLM judge에 의존한다는 점은 향후 보완이 필요하다.