Technical Report for AI4Math-2026 Track 2: Mechanically Orchestrated LLM Sub-Agents for Theorem Proving in Lean 4

저자: Byung-Hak Hwang, Joonhyun La, Chul-hee Lee, Hyojae Lim | 날짜: 2026 | URL: https://openreview.net/forum?id=AJYBzF6wBn 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Pipeline workflow for a single challenge. Five of six sub-agents are shown; Campaign (the outer worklist drive

LLM 기반 Lean 4 정리 증명을 위해 6개의 sub-agent(Generator, Refuter, Formalizer Coordinator, Formalizer Node, Inliner, Campaign)를 deterministic shell state machine으로 조율하는 파이프라인을 제안하고, ICML 2026 AI4Math Challenge 2의 34개 문제 중 30개(600/1000점)를 human-in-the-loop 없이 kernel-verified로 해결했다.

Motivation

Achievement

Figure 1

Figure 1. Pipeline workflow for a single challenge. Five of six sub-agents are shown; Campaign (the outer worklist drive

  1. 파이프라인 설계: Generator, Refuter, Formalizer Coordinator, Formalizer Node, Inliner, Campaign의 6개 sub-agent와 이를 조율하는 deterministic shell state machine을 통해 모든 coordination decision을 reproducible하고 auditable하게 만들었다.
  2. verbatim-preservation 원칙: sub-agent 간 진단 정보(diagnostic output)를 해석 계층 없이 byte-identical하게 전달하는 원칙을 확립했다.
  3. benchmark 성과: ICML 2026 AI4Math Challenge 2의 34개 challenge 중 30개(600/1000점)를 kernel-verified 방식으로 해결했으며, human-in-the-loop 개입 없이 모든 sub-agent dispatch 및 coordination decision을 수행했다.
  4. failure mode 분석: 파이프라인이 겪은 4가지 실패 모드와 그 해결책을 카탈로그화했다.

How

Figure 1

Figure 1. Pipeline workflow for a single challenge. Five of six sub-agents are shown; Campaign (the outer worklist drive

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: LLM 기반 자율 판단을 최소화하고 결정론적 조율 메커니즘으로 대체한다는 아이디어는 formal proof 자동화에서 신뢰성과 재현성을 높이는 실용적이고 참신한 접근이며, 실제 competitive benchmark에서 높은 solve rate를 보여 기술적 타당성을 입증했다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'Minif2f: a cross-system benchmark for formal olympiad-level mathematics'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.94로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'Fimo: A challenge formal dataset for automated theorem proving'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.94로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'Towards large language models as copilots for theorem proving in lean'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구동일한 AI4Math 대회 Track의 후속 파이프라인을 제시한다.
다른 접근동일 대회의 Track 1 논문으로 의미 정합성 검증을 다루며 Track 2와 상보적이다.
기반 연구동일 TCS 증명 벤치마크 시리즈의 다른 Phase 또는 확장 버전을 다루는 밀접한 후속 연구이다.
후속 연구LLM 기반 정리 증명 에이전트 구조를 확장한 연구이다.
다른 접근다중 에이전트 오케스트레이션의 다른 설계를 제시한다.
응용 사례동일한 AI4Math 대회에 파이프라인을 적용한 사례이다.
응용 사례state machine 기반 증명 캠페인의 실제 응용 사례
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드