저자: Haonan Huang | 날짜: 2026 | URL: https://openreview.net/forum?id=5gcshcpnl6 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Figure 1. Grounded scrutiny at scale: architecture, calibration, and execution-dependence. (a) Two-level pipeline: a Pyt
단일 Claude Opus 4.6 에이전트가 published computational physics 논문을 처음부터 재현(reproduction)하는 과정에서, 명시적으로 비판을 요청받지 않았음에도 실행(execution) 기반의 실질적 방법론적 문제(critique)를 자발적으로 제기함을 대규모(111편 Quantum ESPRESSO 논문)와 심층(1편 Nature Communications 논문) 두 스코프에서 보인다.
Figure 2. The Reproduce–Review–Reflect pipeline applied to Pizzi 2016. (a) Three-stage flow: Reproduce (human–agent veri
Figure 3. Review ↔referee overlap and Reflect-stage refinement. Top: sparse 14 × 21 overlap matrix, rows = Review concer
총평: Autonomous LLM agent가 실제 computational physics 논문을 first-principles 계산으로 재현하며 실행 기반의 실질적 critique를 자발적으로 생성한다는 것을 대규모·심층 두 스코프에서 설득력 있게 보여준 흥미로운 연구이며, AI-for-Science 및 자동화된 peer review/재현성 검증의 미래 방향에 중요한 시사점을 제공한다.