⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.
라이선스: OpenReview 공개(오픈액세스)
Essence
Figure 1. Overview of the CLAMP pipeline. A single optimization loop fits one shared, TF-restricted adjacency A. (a) For
CLAMP은 single-cell foundation model이 perturbation prediction에 필요한 regulatory structure를 실제로 보존하는지 여부를 예측 정확도가 아니라 명시적 gene regulatory network 복원을 통해 기계론적으로(mechanistically) 평가하는 프레임워크이다. CRISPR perturbation을 Hill-kinetics ODE 상의 clamp로 취급하고, frozen encoder를 통해 predicted/observed steady-state를 비교함으로써 encoder가 regulatory geometry를 얼마나 잘 보존하는지 측정한다.
Motivation
Known: 기존 연구는 scFM을 fine-tuning하거나 downstream predictor를 통해 perturbation response를 예측하는 방식(GEARS, CPA, AttentionPert 등)으로 평가되어 왔으며, 이러한 regression 기반 벤치마크는 scFM이 additive/mean baseline보다 낮은 성능을 보인다는 사실을 반복적으로 보고해왔다.
Gap: 예측 정확도 기반 평가는 embedding 자체의 정보 손실과 이를 읽는 predictor의 실패를 구분하지 못하는 confound 문제가 있어, scFM이 regulatory structure를 애초에 recoverable한 형태로 보존하는지 여부를 알 수 없다는 근본적 공백이 존재한다.
Why: scFM 사전학습에 투입되는 막대한 자원 대비 실제로 무엇을 학습했는지에 대한 mechanistic 설명이 없다는 점은, 이 모델들이 in-silico perturbation prediction과 target discovery 등 실용적 응용에 신뢰성 있게 사용될 수 있는지를 판단하는 데 핵심적인 문제이다.
Approach: CLAMP은 perturbation을 학습된 ODE 상의 coordinate clamp로 재정의하고, 단일 shared gene regulatory network를 모든 single-gene perturbation 조건에 걸쳐 differentiable dynamical model로 fitting하며, encoder를 교체 가능한 모듈로 두어 recovery quality 자체를 encoder의 regulatory structure 보존 능력에 대한 probe로 활용한다.
Achievement
Figure 3. Structural validation of the recovered adjacency. (A) Combined subnetwork of the top-25 recovered targets per
Identity encoder를 통한 genuine recovery 검증: identity encoder로 복원한 adjacency network가 독립적인 ChIP-seq 및 lineage signature와 정합하며, combinatorial training data 없이 held-out double-gene perturbation(CEBPA+CEBPE 등, ρnonadd=+0.98)을 예측함으로써 복원된 구조가 실제 regulatory structure임을 확인했다.
최초의 scFM별 mechanistic 계층 분석: scGPT, scPRINT, Stack, Tahoe-x1 네 가지 foundation model을 layer 단위로 CLAMP를 통해 읽어, 각 모델이 regulatory structure를 어떻게(어느 층에서 얼마나) 보존하는지에 대한 최초의 기계론적 설명을 제공했다.
Tractable joint fitting 프레임워크 구축: implicit differentiation을 통한 fixed-point solver로 gradient 비용을 solver depth와 분리시켜, Perturb-seq gene-panel 규모에서 단일 shared network의 joint fit을 실현 가능하게 만들었다.
How
Figure 4. Per-layer recovery decomposition and its geometric corroboration for the four foundation-model encoders. Held-
Gene 발현 벡터 x∈R^G에 대해 dxi/dt = bi + Σ Aij h(xj;n,K) − γxi 형태의 Hill-kinetics ODE로 regulation을 모델링.
각 CRISPR perturbation을 해당 유전자를 관측값에 고정(clamp)하는 제약으로 취급하고, 나머지 유전자는 이 clamp 하에서 steady state로 relax.
모든 single-gene 조건에 걸쳐 하나의 shared adjacency matrix A를 fitting하며, predicted와 observed steady state를 frozen encoder(Identity, PCA, scGPT, scPRINT, Stack, Tahoe-x1) 임베딩 공간에서 shift 기반 loss로 비교.
총평: 예측 정확도라는 confound된 척도를 우회하여 scFM의 regulatory structure 보존 여부를 기계론적으로 검증하는 참신하고 방법론적으로 탄탄한 접근이며, single-cell foundation model 평가 방법론에 중요한 새 방향을 제시한다.
기반 연구SPECTER2 유사도 0.92로 Scientific Machine Learning for Dynamics와 Agentic AI for Scientific Automation가 맞닿아, 'CRISPR-GPT for agentic automation of gene-editing experiments'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Scientific Machine Learning for Dynamics와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Efficient fine-tuning of single-cell foundation models enables zero-shot molecular perturbation prediction'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Scientific Machine Learning for Dynamics와 AI-Driven Drug and Materials Discovery가 맞닿아, 'PerTurboAgent: A Self-Planning Agent for Boosting Sequential Perturb-seq Experiments'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.