MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation Effects

저자: Ruihan Guo, Chaoran Cheng, Zhanghan Ni, Neil He, Bangji Yang, Ge Liu | 날짜: 2026 | URL: https://openreview.net/forum?id=ZcVSBTJj2m 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Overview of mutation signal construction and representation. For a residue selected from an experimentally res

단백질 구조 데이터베이스(PDB) 전역에 걸쳐 물리 기반 에너지 모델, 단백질 언어 모델, inverse folding 모델의 단일 부위 돌연변이 신호를 정렬한 대규모 데이터셋 MutAtlas를 구축하고, 이를 바탕으로 실험 라벨 없이 다중 소스 간 불일치를 명시적으로 모델링하는 unsupervised mutation preference distillation 프레임워크를 제안한다.

Motivation

Achievement

Figure 5

Figure 5. Cross-method agreement and ranking structure. Left: Top-1 agreement rates. Right: per-position Spearman rank c

  1. PDB-wide mutation augmentation dataset 구축: 86,479개 protein chain, 15,900,903개 residue position, 3억 건 이상의 single-site mutation evaluation을 포함하는 데이터셋을 구축하여, 기존 최대 DMS 컬렉션 대비 구조적 커버리지가 약 두 자릿수(2 orders of magnitude) 더 크다.
  2. 대규모 cross-source 불일치 분석: unified position-wise mutation preference representation 하에서 FoldX, protein language model, inverse folding model 간 mutation distribution의 consistency, concentration, substitution pattern에 상당한 차이가 있으며 이는 random noise가 아닌 conflicting inductive bias를 반영함을 실증적으로 규명하였다.
  3. unsupervised multi-source distillation 프레임워크 제안: 실험 라벨을 전혀 사용하지 않고도 ProteinGym에서 평가된 zero-shot baseline 및 naive multi-source fusion 전략들보다 우수한 전반적 성능을 달성하였다.
  4. 재현 가능한 자원 공개: 데이터셋, 코드, 평가 파이프라인을 공개하여 커뮤니티의 후속 연구를 지원한다.

How

Figure 1

Figure 1. Overview of mutation signal construction and representation. For a residue selected from an experimentally res

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: PDB 전역을 아우르는 대규모 mutation augmentation 데이터셋과 소스 간 불일치를 명시적으로 다루는 unsupervised distillation 프레임워크를 함께 제시하여 데이터 자원과 방법론 모두에서 실질적 기여를 하는 견고한 연구이다. 다만 세부 방법론의 이론적 근거와 다양한 다운스트림 응용에 대한 검증이 추가되면 더욱 완성도가 높아질 것이다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.92로 Computational Molecular Design와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'Evolutionary-scale prediction of atomic-level protein structure with a language model'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92 기준으로 'MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation Effects'의 AI4S 방법론을 'De novo design of protein structure and function with RFdiffusion'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구SPECTER2 유사도 0.92 기준으로 'MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation Effects'의 AI4S 방법론을 'Agentic End-to-End De Novo Protein Design for Tailored Dynamics Using a Language Diffusion Model'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구PLM 표현을 보완하는 multi-modal 구조를 확장한 연구로 판단된다.
응용 사례단백질 언어 모델과 inverse folding 신호를 활용하는 유사한 응용 사례임
기반 연구단백질 언어 모델의 기초적인 표현 학습 방법론을 제공한다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Bi-TEAM: A Unified Cross-Scale Representation Learning Framework for Chemically Modified Biomolecules'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근단백질-화학 정보 통합을 위한 다른 접근법을 제시
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'ProMaya: a hierarchical universal Deep Learning framework for accurate and interpretable Protein-Protein interaction identification'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근단백질 돌연변이 신호를 다루는 다른 데이터셋 구축 및 평가 방법을 제시함
다른 접근다중 소스 신호 정렬이라는 유사 문제를 다른 방식으로 접근한다.
후속 연구protein language model의 latent 표현을 활용한 epistasis 분석의 이론적 기반이 될 수 있음
응용 사례물리 기반 에너지 모델을 실제 단백질 데이터에 적용한 사례를 다룬다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드