Latent Encoding of Epistasis in Protein Language Models

저자: Kun Hyung Roh, Vincent Yip, Sunghoon Rho | 날짜: 2026 | URL: https://openreview.net/forum?id=Dc3xP3f3Fb 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Monotone response calibration for measured stability.

단백질 다중 돌연변이 landscape를 additive 효과, monotone global response, specific epistasis residual의 3계층으로 분리하고, frozen diffusion language model 표현에 학습된 경량 pairwise head를 얹어 이 residual coupling 신호를 예측하는 프레임워크를 제안한다.

Motivation

Achievement

Figure 3

Figure 3. (A) The coupling head beats the zero shot score on all eleven measured function assays. (B) Ablations on GB1 l

  1. 3계층 분해 프레임워크 정립: 측정된 multi-mutant 데이터를 additive, monotone global, specific epistasis residual로 분리하는 수학적 틀(Equations 1-8)을 제시하고, 이 residual이 additive 모델이나 additive score의 monotone 변환으로는 원리적으로 접근 불가능함을 증명했다.
  2. 대규모 실증 평가: 149개 protein domain에 걸친 89,706개 stability double mutant와 11개 functional assay에 걸친 85,135개 multi-mutant 데이터에서 zero-shot DPLM 점수가 stability coupling은 제한적으로, functional coupling은 거의 포착하지 못함을 확인했다.
  3. 학습된 coupling head의 성능 향상: 학습된 representation 기반 모델이 11개 assay 전체에서 functional epistasis를 회복하며, median held-out Spearman이 0.008에서 0.292로, GB1 binding에서는 0.02에서 0.37로 크게 개선되었다.
  4. 실용적 설계 활용 검증: additive-matched candidate pool(additive 선택이 무의미한 구간)에서도 학습된 residual 모델이 gain-of-function double mutant를 성공적으로 선별함을 GB1 landscape에서 입증했다.

How

Figure 3

Figure 3. (A) The coupling head beats the zero shot score on all eleven measured function assays. (B) Ablations on GB1 l

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Specific epistasis를 additive/global 성분과 엄밀히 분리해 측정 가능한 학습 목표로 재정의한 개념적 기여가 명확하고, 다수의 실제 measured landscape에서 이를 실증적으로 뒷받침한 견고한 연구이나, wet-lab 검증과 higher-order epistasis로의 확장은 향후 과제로 남아있다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.90로 Computational Molecular Design와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구돌연변이 효과 분해를 위한 이론적 기초를 제공하는 연구로 판단됨
기반 연구protein language model의 latent 표현을 활용한 epistasis 분석의 이론적 기반이 될 수 있음
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Computational Molecular Design와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'How to make the most of your masked language model for protein engineering'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Linear-time prediction of proteome-scale microbial protein interactions'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.91로 Computational Molecular Design와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'AROMA: Augmented Reasoning Over a Multimodal Architecture for Virtual Cell Genetic Perturbation Modeling'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Biologically-Grounded Multi-Encoder Architectures as Developability Oracles for Antibody Design'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근단백질 epistasis 표현을 위한 다른 protein language model 활용 방식으로 보임
후속 연구단백질 fitness landscape 분해 방법론을 확장한 후속 연구로 보임
후속 연구frozen language model 표현을 활용한 관련 확장 연구로 보임
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드