Tacit Verification Apprenticeship for Self-Evolving Scientific Agents

저자: David Scott Lewis, Karl Wang, Zhaoxiang Feng, Junhan Wang, Saien Deng | 날짜: 2026 | URL: https://openreview.net/forum?id=cxNZmu7ODQ 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. The TVA loop. Verification failure is treated as training

이 논문은 formal verification 도구(증명 보조기, 검증기, 모델 체커)를 호출할 수 있는 능력만으로는 self-evolving scientific agent가 신뢰할 수 없다고 주장하며, 정의 선택, 증명 의무 분할, 라이브러리 탐색, 실패한 스크립트 수리, 형식적으로는 accepted되었지만 과학적으로 misaligned된 artifact를 판별하는 능력인 "tacit verification craft"를 학습하는 아키텍처 TVA(Tacit Verification Apprenticeship)를 제안한다.

Motivation

Achievement

Figure 3

Figure 3. Experiment 1: mean craft-similarity profiles for the dedu-

  1. Tacit Verification Bottleneck 정의: formal artifact와 이를 만들고 진단·수리·전이하는 데 필요한 craft 사이의 간극을 공식적으로 개념화했다.
  2. TVA 아키텍처 제안: 다섯 개 모듈로 구성된 apprenticeship 기반 architecture와 typed memory schema, skill utility 함수(수식 1-3), failure diagnosis, repair selection, transfer gating 알고리즘을 제시했다.
  3. ApprenticeshipBench-SESA 벤치마크 설계: formal math, semantic alignment, proof repair, executable scientific spec, domain validation을 아우르는 평가 체계를 제안했다.
  4. 경험적 probe 결과: craft 차원이 corpus에서 recoverable함을 보였고, apprenticeship retrieval이 process annotation으로 개선되며, repair-memory transfer는 명시적 gate가 필요하고, sentinel이 semantic disagreement를 repair나 abstention으로 전환할 수 있음을 세 가지 corpus probe와 semantic-sentinel demo로 입증했다.

How

Figure 2

Figure 2. Deduplicated top-cluster corpus used for the experiments.

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 3/5 Significance: 4/5 Clarity: 4/5 Overall: 3/5

총평: Formal verification 도구를 다루는 데 필요한 암묵적 전문가 craft를 학습 대상으로 정식화한 개념적 기여가 인상적이지만, 이를 뒷받침하는 실증 결과는 소규모 corpus probe 수준에 머물러 있어 실제 대규모 self-evolving scientific agent 시스템으로서의 완성도는 향후 연구를 통해 더 검증되어야 한다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 Scientific Information Extraction and QA가 맞닿아, 'Fact-checking complex claims with program-guided reasoning'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Autoreproduce: Automatic AI Experiment Reproduction with Paper Lineage'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'ResearchGym: Evaluating Language Model Agents on Real-World AI Research'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근증명 보조기 및 검증기 기반 agent의 신뢰성 문제를 다루는 유사 접근
다른 접근과학적 tacit knowledge 습득이라는 유사한 문제의식을 공유하는 연구
다른 접근formal verification 도구를 활용한 자기진화 에이전트의 대안적 설계
후속 연구merge-readiness 판단을 위한 벤치마크 설계의 기초적 방법론을 제공한다.
응용 사례과학적 에이전트에 형식 검증을 적용한 사례
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드