CertJudge: Evaluating Lean Formal-Code With Falsifiable Properties

저자: Ethan S Hersch, Brando Miranda, Elyas Obbad, Srivatsava Daruru, Zhanke Zhou, Kirill Acharya, Sanmi Koyejo | 날짜: 2026 | URL: https://openreview.net/forum?id=xmHP3UhGQ8 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Scalable trust calibration for LLM judges. A small

LLM judge가 Lean 4 형식 정형 산출물(정리/명세)의 품질을 평가할 때, 매번 새로운 human validation 없이도 신뢰도를 자동으로 검증할 수 있는 property-certification 프로토콜인 CertJudge를 제안한다. identity, bug monotonicity, specification monotonicity, stability의 네 가지 falsifiable property를 조합한 variance-aware trust index (TI_var)를 통해 judge의 human alignment를 예측한다.

Motivation

Achievement

Figure 3

Figure 3. TIvar vs. judge-level human alignment. Each point is one judge configuration; the x-axis is TIvar, and the y-a

  1. amortized trust-validation protocol 제안: 고정된 Lean artifact panel에 대한 perturbation 기반 property certification을 일회성 calibration으로 정식화하여, 이후 새로운 judge 구성은 human labeling 없이 네 가지 falsifiable property만 계산해 스크리닝할 수 있게 했다.
  2. n=75 규모의 proof-of-concept calibration 수행: TI_var가 진단들의 geometric aggregation으로서 human agreement를 잘 추적함을 검증셋에서 Spearman correlation 0.833(avg labels), 0.905(pass1), 0.738(pass2)로 보여주었다.
  3. VeriBench에서의 end-to-end 활용: calibrated judge로 theorem-generation 방법들을 순위화하여(예: DSPy ReAct 0.615 vs baseline prompting 0.470, normalized specs 기준) benchmark 기대치와 일치하는 결과를 도출, 소규모 human calibration set이 scalable judge selection을 지원할 수 있음을 실증했다.

How

Figure 4

Figure 4. (P3) Monotonicity w.r.t. missing specifications. Candidate scores decrease as more theorems/tests are withheld

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 3/5 Significance: 4/5 Clarity: 4/5 Overall: 3/5

총평: scalable oversight라는 중요한 문제에 대해 falsifiable property 기반의 실용적이고 창의적인 접근을 제시했으나, calibration 규모가 작고 generalization 검증이 부족해 proof-of-concept 단계에 머무른다는 점에서 추가 검증이 필요한 초기 연구이다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.92로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Autoreproduce: Automatic AI Experiment Reproduction with Paper Lineage'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'Grammars of formal uncertainty: When to trust llms in automated reasoning tasks'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구verifier 기반 필터링에서의 mode collapse 현상을 확장하여 분석한다.
기반 연구형식 검증 산출물의 품질 평가에 필요한 property 기반 검증 방법론의 기초를 제공한다.
기반 연구SPECTER2 유사도 0.92로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'M2F: Automated Formalization of Mathematical Literature at Scale'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근다른 도메인의 형식적 산출물에 대해 유사한 자동 인증 접근을 취한다.
후속 연구Lean 기반 형식 검증 아티팩트 생성의 방법론적 기초를 제공하는 연구로 보임
다른 접근identity/property 기반 검증이라는 유사한 인증 방법론
다른 접근LLM judge의 신뢰도를 인간 검증 없이 자동으로 확보하는 유사한 평가 프로토콜을 제안한다.
다른 접근Lean 형식 검증 관련 유사한 자동화 검증 접근법을 제시한다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드