MADS-CPS: A Machine-Checkable Admissibility Contract for AI Scientists in Autonomous Laboratories

저자: Mateo PETEL | 날짜: 2026 | URL: https://openreview.net/forum?id=VrYHFXGyUO 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. MADS-CPS evaluates each run relative to a declared assurance envelope and emitted artifact family. Admissibili

AI scientist 시스템이 자율 실험실에서 실제 물리적 실행을 수행하게 되면서, 능력(capability) 평가만으로는 부족하고 한 번의 구체적 run이 감사 가능하고(inspectable) 재현 가능하며(replayable) 무결성을 유지하고(integrity-preserving) 비가역적 지점(PONR)에서 통제되는지를 기계적으로 검증하는 run-level admissibility contract인 MADS-CPS를 제안한다.

Motivation

Achievement

Figure 1

Figure 1. MADS-CPS evaluates each run relative to a declared assurance envelope and emitted artifact family. Admissibili

  1. run-level admissibility contract 정의: declared envelope, required artifact(구조화된 실행 trace, 평가 보고서, evidence bundle, release manifest), admissibility predicate, conformance tier, verification mode, PONR release semantics로 구성된 계약을 최초로 명세.
  2. controller-agnostic conformance 공식화: assurance layer에서 checker가 controller 내부 로직이 아니라 방출된 evidence만을 대상으로 동작하도록 형식화.
  3. restricted-auditability 모델 제시: payload 내용이 redaction 되었을 때 어떤 predicate가 손실되는지 명시하면서도 구조적 무결성(structural integrity)을 보존.
  4. 재현 가능한 reference instantiation과 4종 평가: eight-case conformance challenge corpus, verification-mode admissibility matrix, independent replay-link experiment, controller-matrix study(baseline vs stressed regime)를 통해 checker가 모든 injected evidence-fault label과 일치하고, independent replay가 producer report와 일치하며, coordination stress 하에서 evidence-layer admissibility는 안정적으로 유지되는 반면 operational productivity는 divergence를 보임을 입증.

How

Figure 1

Figure 1. MADS-CPS evaluates each run relative to a declared assurance envelope and emitted artifact family. Admissibili

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 3/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: AI scientist의 물리적 실행이 늘어나는 시점에 capability와 admissibility를 명확히 분리하는 run-level evidence contract를 제시한 시의적절하고 개념적으로 탄탄한 연구이나, 제공된 발췌만으로는 실험적 검증의 규모와 실제 실험실 적용 가능성에 대한 확신을 얻기는 제한적이다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.93로 Formal Proof Verification Automation와 Agentic AI for Scientific Automation가 맞닿아, 'AIGS: Generating science from ai-powered automated falsification'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 Formal Proof Verification Automation와 Agentic AI for Scientific Automation가 맞닿아, 'SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구AI 시스템의 검증 가능성에 대한 형식적 계약 기반을 제공한다.
기반 연구자율 실험실 AI 시스템의 무결성 보장을 위한 이론적 토대를 제공한다.
기반 연구SPECTER2 유사도 0.93로 Formal Proof Verification Automation와 Agentic AI for Scientific Automation가 맞닿아, 'ENPIRE: Agentic Robot Policy Self-Improvement in the Real World'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근AI 과학자 실행에 대한 다른 감사·재현성 프레임워크를 다룬다.
다른 접근AI scientist 시스템의 안전성/감사 가능성을 다루는 유사한 프레임워크이다.
응용 사례자율 실험실 환경에 무결성 검증 계약을 실제 적용한다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드