Essence
Figure 1. Lower-ladder outcomes on standalone_mix100.
Agentic Lean prover의 pass rate가 도구 사용, 검색, 검증 등 다양한 요소가 얽혀 해석이 어렵다는 문제를 지적하고, 100-task 고정 벤치마크에서 trace-level ablation을 통해 어떤 요소(compiler feedback, proof-state feedback, tool access, retrieval)가 실제로 성능에 기여하는지 단계별로 분석한다.
Evaluation
Novelty: 4/5 Technical Soundness: 3/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5
총평: Agentic Lean prover의 성능을 구성 요소별로 분해하여 분석한 실용적이고 시의적절한 attribution 연구로, 특히 tool access와 tool use의 구분 및 commitment failure 발견은 향후 agent 설계에 중요한 시사점을 준다. 다만 소규모 벤치마크와 mixed된 통계적 결과로 인해 일부 결론의 확증성은 제한적이며 workshop-level 초기 연구로서의 성격이 강하다.
같이 보면 좋은 논문
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Autoreproduce: Automatic AI Experiment Reproduction with Paper Lineage'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'LLM Agents Making Agent Tools'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구LLM 에이전트의 end-to-end workflow 평가라는 유사한 프레임워크를 확장함
기반 연구agentic prover의 구성요소별 성능 분석의 기초가 되는 연구
기반 연구trace-level ablation 분석의 방법론적 기초 제공
기반 연구closed proof 실패 인증 개념을 확장한 연구
기반 연구proof refactoring을 위한 controllable optimization을 확장한다.
다른 접근agentic lean prover의 성능 요인을 다른 방식으로 분석하는 연구
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'M2F: Automated Formalization of Mathematical Literature at Scale'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근Lean proof의 모듈화 및 재구성을 다루는 관련 연구임.
다른 접근도구 오케스트레이션과 LLM 추론 통합을 위한 대안적 시스템 구조 제시
후속 연구trace-level attribution 분석을 확장한 관련 연구