Essence
Figure 1. (a) Execution-grouned evaluation uncovers failures that narrative-alone review misses. In this example, Failur
논문 내러티브만으로는 감지할 수 없는 연구의 문제점을 발견하기 위해, 코드와 데이터를 함께 검증하는 execution-grounded evaluation 프레임워크를 제안하고 MechEvalAgent를 구현했다.
Evaluation
Novelty: 4/5 Technical Soundness: 3/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5
총평: 이 논문은 재현성 위기 시대에 AI 에이전트를 평가자로 활용하는 혁신적 접근을 제시하며, execution-grounded evaluation으로 인간 리뷰어가 놓치는 51개의 문제를 식별하여 과학적 엄밀성 강화의 실질적 경로를 제시한다.
같이 보면 좋은 논문
기반 연구REFORMS의 ML 기반 과학 연구 재현성 체크리스트는 코드와 데이터를 함께 검증하는 execution-grounded evaluation 프레임워크와 결합하면 더 완전한 검증 체계를 구성한다.
기반 연구실제 논문 작업에 AI 에이전트를 적용해 평가한 유사 사례이다.
다른 접근논문 검증을 위한 다른 자동화된 평가 프레임워크를 제안한다.
기반 연구코드 및 실험 재현성 검증의 방법론적 기반을 제공한다.
후속 연구SPECTER2 유사도 0.96 기준으로 'ARA: Agent-Native Research Artifacts'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구완전한 데이터 출처 정보를 포함하는 LLM-native figures는 코드와 데이터를 함께 검증하는 execution-grounded evaluation의 시각적 증거를 제공하는 인터페이스가 될 수 있다.
다른 접근연구 결과 검증이라는 동일 목표를 다른 방법론(execution-grounded vs agentic metrics)으로 다룬다.
다른 접근코드/데이터 실행 검증과 워크플로우 그래프 재현성 평가라는 유사하지만 다른 접근을 제시한다.
다른 접근과학적 주장의 타당성을 검증하는 유사한 목적의 시스템이다.
후속 연구SPECTER2 유사도 0.96 기준으로 'FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
다른 접근연구 루프 전반을 평가하는 벤치마크로서 방법론적 연관성이 있다.
다른 접근코드와 데이터 실행 기반 검증이라는 다른 방식으로 연구 신뢰성을 확보하는 접근이다.
응용 사례실행 기반 검증 방법이 SCION 같은 agentic 연구 실행 시스템의 신뢰성 평가에 활용될 수 있다.
후속 연구SPECTER2 유사도 0.95 기준으로 'Position: Correct Answer, Wrong Mechanism - When AI Scientists Defend General Claims Their Own Data Contradicts'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
후속 연구SPECTER2 유사도 0.95 기준으로 'Reasoning Is More Than the Model: Harness-Aware Evaluation of Agents on Verifiable Reasoning Tasks'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
후속 연구SPECTER2 유사도 0.92 기준으로 'RSI for Science: A Verifier-First Framework for AI Scientists'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
후속 연구SPECTER2 유사도 0.92 기준으로 'VeRA: Math Benchmarks as Executable Specifications'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
후속 연구SPECTER2 유사도 0.92 기준으로 'AI Coding Benchmarks Need Proofs, Not Just Tests'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
후속 연구SPECTER2 유사도 0.94 기준으로 'Experimental Attempts in Electronic Lab Notebooks: A Dataset Proposal for Scientific Debugging'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
후속 연구SPECTER2 유사도 0.94 기준으로 'FabScore: Fine-Grained Evaluation of Fabrications in Automated AI Research'의 AI4S 방법론을 'The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.