Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design

저자: Gyubok Lee, Kiwoong Yoo, Jimin Seo, Kyunghoon Hur, Edward Choi | 날짜: 2026 | URL: https://openreview.net/forum?id=7DvYeitu6I 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Overview of the post-generation binder shortlisting task. For each target protein, a fixed pool of candidate b

단백질 binder de novo 설계 파이프라인에서 후보 생성 이후의 병목인 shortlisting 문제를, 사전 계산된 17개 구조적 신뢰도/interface 품질 proxy 점수로부터 LLM이 다중 지표 랭킹 정책(policy)을 생성하도록 하여 해결하는 방법을 제안한다.

Motivation

Achievement

Figure 2

Figure 2. Feature weights in target-conditioned iterative gpt-4o ranking policies over the full 17-feature panel. The ve

  1. 10-target held-out 성능 개선: 5개 샘플링된 global iterative gpt-4o policy 평균이 Recall@10 0.589를 달성하여, 가장 강력한 단일 feature 고정 baseline인 Protenix binder ipTM(0.571)을 소폭 상회했다.
  2. 3-target subset 최고 LLM 성능: Nipah, RBX1, TREM2로 구성된 3-target held-out subset에서 target-conditioned iterative gpt-5.4 policy가 Recall@10 0.519, NDCG@10 0.583으로 LLM 기반 방법 중 최고 성능을 보였다.
  3. 해석 가능한 결정 계층 제시: LLM이 생성한 랭킹 정책이 이질적 구조적 신뢰도 및 interface 품질 proxy 지표를 결합하는 해석 가능한 feature-weighted 조합으로 작동함을 보였다.

How

Figure 1

Figure 1. Overview of the post-generation binder shortlisting task. For each target protein, a fixed pool of candidate b

Originality

Limitation & Further Study

Evaluation

Novelty: 3/5 Technical Soundness: 3/5 Significance: 3/5 Clarity: 4/5 Overall: 3/5

총평: 단백질 binder 설계 파이프라인의 실질적 병목인 shortlisting 문제에 LLM 기반 해석 가능 정책 생성이라는 참신한 응용을 제시했으나, held-out target 수가 적고 성능 향상 폭이 제한적이어서 워크숍 수준의 초기 탐색 연구로 평가된다.

같이 보면 좋은 논문

기반 연구단백질 설계 방법론을 실제 항체 개발에 적용한 사례를 제공함
기반 연구AlphaFold-Multimer hallucination 기법을 확장하여 응용한다.
기반 연구단백질 binder 설계의 구조적 신뢰도 평가 지표에 기초를 제공한다.
후속 연구de novo protein binder design 방법론의 기초가 되는 연구이다.
기반 연구SPECTER2 유사도 0.94로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Co-designing sequence and structure of functional de novo enzymes with EnzyGen2'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 Computational Molecular Design와 AI-Driven Drug and Materials Discovery가 맞닿아, 'InstructNA leverages nucleic acid large language models with HT-SELEX for de novo generation of functional nucleic acids'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93 기준으로 'Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design'의 AI4S 방법론을 'On the Reliability of AI Methods in Drug Discovery: Evaluation of Boltz-2 for Structure and Binding Affinity Prediction'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구SPECTER2 유사도 0.94로 Computational Molecular Design와 Molecular Simulation and Generative Modeling가 맞닿아, 'Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근약물 발견을 위한 모델 성능 비교에서 대안적 벤치마킹 방법론을 제시한다.
기반 연구지속적 평가 체계를 다른 생체분자 예측 과제로 확장한다.
다른 접근후보 shortlisting 문제를 LLM이 아닌 다른 방식으로 해결한다.
후속 연구epitope-conditioned binder 생성의 기초가 되는 단백질 언어모델 연구이다.
후속 연구Function-guided 단백질 생성의 이론적 기초 제공
후속 연구set-level diversity 보상 설계의 기반이 되는 연구이다.
응용 사례다중 지표 기반 랭킹 정책을 실제 binder 설계에 적용한다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드