A Statistical Framework for Mechanistic Claims in Neural Networks: The Predict--Intervene--Validate Pipeline

저자: Zacharie Bugaud | 날짜: 2026 | URL: https://openreview.net/forum?id=o8aRR9YuxE 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

이 논문은 mechanistic interpretability에서 이루어지는 인과적 주장(mechanistic claims)을 통계적으로 검증하기 위한 표준화된 3단계 파이프라인인 PIV (Predict-Intervene-Validate)를 제안하고, 이를 Elman RNN의 input-invariant dimensions 메커니즘에 적용하여 그 판별력을 실증한다.

Motivation

Achievement

  1. Predict 단계 성능: 300개의 fresh held-out baseline Elman RNN에 대해 사전등록된 margin certificate가 AUROC 0.97, precision-optimal operating point에서 97.8% precision(balanced accuracy 0.71)을 달성했으며, 15개 candidate predictor 중 13개가 기각되고 gap formula 등 2개만 확인되었다.
  2. Intervene 단계 성능: confound-controlled weight surgery(랜덤 orthonormal 방향 및 matched-rank perturbation을 대조군으로 사용)에서 393/393 인과적 손상(causal break)이 관찰되었고, cluster-bootstrap 검정으로 p<10^-30의 유의성을 얻어 within-model dependence에 강건함을 보였다.
  3. Validate 단계 성능: hyperparameter를 고정한 채 12.5배 스케일링된 OOD 전이 실험에서 160/160의 성공률을 보였으며, model-level over-dispersion을 고려한 beta-binomial null model 하에서 p<10^-71의 극도로 유의한 결과를 얻었다.
  4. 실용적 산출물: 각 PIV 단계에 대한 sample-size 가이드라인, mechanistic interpretability 제출 논문을 위한 체크리스트, piv Python package, 데모 노트북을 공개하여 실무 적용성을 높였다.

How

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: mechanistic interpretability 분야의 오랜 방법론적 공백을 통계학의 표준 도구(사전등록, power analysis, confound control)로 메우려는 시도로서 실용적 가치가 높고, 구체적 사례연구를 통한 실증도 설득력 있으나, 단일 사례에 국한된 검증이라 더 넓은 아키텍처·현상으로의 일반화 검증이 후속 과제로 남아있다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.93로 Statistical Causal Inference Methods와 Agentic AI for Scientific Automation가 맞닿아, 'Biodsa-1k: Benchmarking data science agents for biomedical research'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구생물학적 데이터에 다중 ML 기법을 적용하는 유사한 방법론적 접근을 공유한다.
기반 연구인과적 주장 검증의 통계적 기초를 제공하는 관련 연구이다.
기반 연구concept steering의 representation 검증을 확장하는 연구
기반 연구SPECTER2 유사도 0.92로 Statistical Causal Inference Methods와 Agentic AI for Scientific Automation가 맞닿아, 'Scaling Reproducibility: An AI-Assisted Workflow for Large-Scale Reanalysis'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Statistical Causal Inference Methods와 Agentic AI for Scientific Automation가 맞닿아, 'Celcomen: spatial causal disentanglement for single-cell and tissue perturbation modeling'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근mechanistic interpretability의 인과적 검증이라는 동일 문제를 다른 통계적 틀로 접근한다.
다른 접근interpretability claim의 실증적 검증 절차를 제안하는 유사한 접근
응용 사례신경망 메커니즘 해석에 통계적 검증 방법을 적용한 사례이다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드