E1: Retrieval-Augmented Protein Encoder Models

저자: Sarthak Jain, Joel Beazer, Jeffrey A. Ruffolo, Aadyot Bhatnagar, Ali Madani | 날짜: 2026 | URL: https://openreview.net/forum?id=cKrZo7L60A 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. E1 Architecture. The E1 model can take in homologous sequences in addition to an input query sequence. The hom

E1은 homologous sequence를 block-causal multi-sequence attention으로 명시적으로 조건화하는 retrieval-augmented protein language model 계열로, fine-tuning 없이 zero-shot fitness 예측과 unsupervised contact-map 예측에서 state-of-the-art 성능을 달성한다.

Motivation

Achievement

Figure 2

Figure 2. Examples from CAMEO dataset where retrieval augmentation helps E1 identify contact it may have mispredicted wh

  1. State-of-the-art zero-shot fitness 예측: ProteinGym 217개 Deep Mutational Scan assay에서 E1이 single-sequence 모드로도 ESM 계열을 능가하고, retrieval-augmented 모드에서는 PoET, MSA Pairformer 등 기존 retrieval 기반 모델도 능가함.
  2. Unsupervised contact-map 예측 우수성: single-sequence 모드에서 ESM 계열보다 우수하며, retrieval 추가 시 추가적인 큰 성능 향상을 보임.
  3. Parameter scaling 검증: 150M부터 600M 파라미터까지 모델 크기에 따라 성능이 일관되게 향상됨을 확인.
  4. 유연한 추론 모드: 동일 모델이 single-sequence 또는 retrieval-augmented 모드로 fine-tuning 없이 전환 가능하여 fitness prediction, variant ranking, 구조 관련 embedding 생성에 모두 활용 가능.
  5. 오픈 소스 공개: 150M/300M/600M 세 가지 E1 모델을 연구 및 상업적 사용을 위해 무료로 공개.

How

Figure 1

Figure 1. E1 Architecture. The E1 model can take in homologous sequences in addition to an input query sequence. The hom

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Alignment-free retrieval augmentation과 효율적인 alternating block-causal attention을 결합해 single-sequence 및 retrieval 모드 모두에서 SOTA를 달성하고 모델을 오픈소스로 공개했다는 점에서 실용적 가치와 학술적 기여도가 높은 연구다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'Evolutionary-scale prediction of atomic-level protein structure with a language model'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구key-aware retrieval 기법을 의료 도메인에 확장 적용함
기반 연구STRING 데이터베이스 기반 PPI 예측을 실제 도메인에 적용한 사례
기반 연구interaction 예측을 위한 attention 기반 접근을 확장한 연구임.
기반 연구retrieval 기반 단백질 표현 학습의 이론적 기반을 공유한다.
후속 연구protein design의 기초가 되는 language model 연구
기반 연구사전학습된 BioFM을 다중 modality 문제에 적용한다.
기반 연구SPECTER2 유사도 0.93 기준으로 'E1: Retrieval-Augmented Protein Encoder Models'의 AI4S 방법론을 'An equivariant pretrained transformer for unified 3D molecular representation learning'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'Protein Language Models Diverge from Natural Language: Comparative Analysis and Improved Inference'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.94로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Linear-time prediction of proteome-scale microbial protein interactions'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'ProMaya: a hierarchical universal Deep Learning framework for accurate and interpretable Protein-Protein interaction identification'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근protein language model의 zero-shot 성능 향상을 위한 다른 아키텍처 접근을 제시한다.
후속 연구단백질 서열 기반 검색의 공통 기초 방법론을 다룬다.
다른 접근protein language model과 구조 정보를 결합하는 유사한 항체 설계 방법이다.
후속 연구단백질 임베딩 기반 검색이라는 유사한 기초 방법론을 공유한다.
후속 연구block-causal attention 메커니즘을 확장하여 적용한 후속 연구이다.
응용 사례retrieval-augmented protein 모델을 실제 단백질 기능 예측에 적용한 사례이다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드