Semi-knockoffs: a model-agnostic conditional independence testing method with finite-sample guarantees

저자: Angel David REYERO LOBO, Bertrand Thirion, Pierre Neuvial | 날짜: 2026 | URL: https://openreview.net/forum?id=Xf9hJMGwDd 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Optimization stability. Data are generated from z =

Semi-knockoffs는 사전학습된 임의의 ML 모델을 그대로 활용하면서도 train-test split 없이 유한 표본에서 유효한 p-value와 FDR 제어를 제공하는 model-agnostic conditional independence testing(CIT) 방법이다. model-X 가정(입력 분포에 대한 완전한 지식) 대신 연속 변수에 대한 conditional expectation 추정만을 요구하여 실용성을 크게 높인다.

Motivation

Achievement

Figure 6

Figure 6. Discoveries on real data across ML models. The

  1. Semi-knockoffs 절차 제안: 임의의 사전학습된 ML 모델을 받아들이면서도 train-test split과 exact knockoff variable 구성 없이, knockoff statistic에 부과되는 제약도 회피하는 새로운 CIT 방법을 제시하고, finite-sample type-I error 및 FDR control을 증명하였다.
  2. Estimated sampler 하에서의 이론적 결과: (i) null feature로 학습된 regularized 모델의 stability를 증명하여 Semi-knockoffs의 distributional convergence를 보이고, (ii) double-robustness 성질을 제시하여 모델 또는 sampler 중 하나만 정확해도 제어가 성립함을 규명하였다(모델과 sampler가 함께 빠르게 수렴할 때의 제어를 conjecture로 제시).
  3. 광범위한 실험적 검증: 최신 variable selection/importance 방법들과 비교하여 Optimization stability(Fig 1), Double Robustness(Fig 3), adjacent support 및 masked correlation 하의 Type-I error(Fig 4, 5), 그리고 실제 데이터에서 여러 ML 모델에 걸친 discoveries(Fig 6)를 통해 방법의 유효성과 실용성을 입증하였다.

How

Figure 3

Figure 3. Empirical evidence for Double Robustness: Dis-

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Model-X 가정을 완화하면서 train-test split 없이 finite-sample 통계 보장을 제공하는 실용적이고 이론적으로 탄탄한 CIT 방법으로, ML 모델을 활용한 과학적 발견 분야에 실질적 기여를 하는 논문이다. 다만 double-robustness의 핵심 결과가 conjecture 수준에 머무는 점은 후속 연구로 보완되어야 한다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.91 기준으로 'Semi-knockoffs: a model-agnostic conditional independence testing method with finite-sample guarantees'의 AI4S 방법론을 'REFORMS: Consensus-based Recommendations for Machine-learning-based Science'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구knockoff 기반 FDR 제어 이론의 기초를 제공한다
기반 연구커널 기반 이표본 검정의 검정력 최적화라는 공통 주제를 확장하여 다룬다.
기반 연구SPECTER2 유사도 0.90로 Statistical Causal Inference Methods와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Representative, Informative, and De-Amplifying: Requirements for Robust Bayesian Active Learning under Model Misspecification'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.90로 Statistical Causal Inference Methods와 Agentic AI for Scientific Automation가 맞닿아, 'Celcomen: spatial causal disentanglement for single-cell and tissue perturbation modeling'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.90로 Statistical Causal Inference Methods와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Knowing when to trust machine-learned interatomic potentials'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근model-agnostic conditional independence testing이라는 동일한 문제를 다른 방식으로 해결한다.
후속 연구e-value 기반 순차 검정과 관련된 kernel 추론 방법론적 기반을 제공한다.
다른 접근두 표본 검정에서 특징 기여도를 다르게 추정하는 대안적 접근을 제시한다.
후속 연구유한 표본에서의 유효성을 보장하는 검정 방법을 확장한다
후속 연구kernel 기반 추론 방법론과 밀접하게 연관된 통계적 기반 연구이다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드