PepLang-Bench - Evaluating Large Language Models Understanding On Peptide Related Tasks

저자: Deepa Mal Korani, Vinay Jethava, Jovan Damjanovic, Solmaz Gabery Adams, Kristine Deibler | 날짜: 2026 | URL: https://openreview.net/forum?id=wE0ngq19Kx 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Overview of PepLang-Bench. (a) Datasets: Two peptide subsets; a therapeutic set composed of naturally occurrin

PepLang-Bench는 property prediction, notation conversion, structural adaptation 세 범주 7개 과제로 구성된 peptide 특화 LLM 벤치마크로, generalist 모델(GPT-5, o4-mini, Gemini-3-Pro, Qwen 3.5)과 domain-specific 화학 모델(ChemDFM-R, ChemDFM-v2.0, Ether-0, NatureLM-8x7B) 총 8개 모델의 peptide 이해·추론 능력을 체계적으로 평가한다.

Motivation

Achievement

Figure 1

Figure 1. Overview of PepLang-Bench. (a) Datasets: Two peptide subsets; a therapeutic set composed of naturally occurrin

  1. 혼재된 generalist 성능 패턴 규명: GPT-5는 GRAVY에서 99.7%지만 pI에서는 21.1%에 그치고, Gemini-3-Pro는 반대로 GRAVY 4.0%, pI 64.0%를 기록하는 등 모델별로 상반된 강약점이 드러났다.
  2. code interpreter의 과제 의존적 효과 발견: 코드 실행이 pI(+70%), instability index(+86%), 대부분의 notation conversion(L→C therapeutic +84%, C→L therapeutic +81%)을 크게 개선하지만, therapeutic peptide의 SMILES→HELM 변환은 93.8%에서 82.5%로 오히려 하락하는 trade-off를 확인했다.
  3. noncanonical residue에 의한 일반화 격차 노출: Aib, D-amino acid 등 non-proteinogenic 잔기를 포함한 synthetic peptide에서 SMILES→HELM 정확도가 GPT-5는 93.8%→26.2%, Gemini-3-Pro는 81.0%→27.0%로 급락하여 현재 화학 특화 사전학습이 다루지 못하는 distribution gap을 드러냈다.
  4. domain-specific model의 체계적 실패 모드 분석: ChemDFM-R 등 domain-specific model은 notation·structural adaptation 과제에서 0.92 이상의 confidence를 보고하면서도 정확도 0%를 기록했으며, amino acid를 주기율표 원소로 오인하는 domain confusion, 반복적 degenerate reasoning loop, sequence hallucination, non-canonical residue 무시 등의 구체적 오류 패턴을 확인했다.
  5. 전문가 prompt 개선 연구: linear-to-cyclic 변환 과제에서 도메인 전문가와 함께 prompt를 개선한 결과 측정 가능한 성능 향상은 있었으나 실용 수준에는 미치지 못해, 더 나은 prompting이 근본적 격차를 해소하지는 못함을 보였다.

How

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Peptide라는 중요하지만 그동안 간과된 화학 모달리티에 대해 최초로 체계적인 LLM 벤치마크를 제시하고, code interpreter의 양면성과 non-canonical residue에 대한 일반화 격차 등 실용적으로 중요한 통찰을 제공하는 의미 있는 workshop-level 연구이다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.92로 Computational Molecular Design와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'Accelerating science with human-aware artificial intelligence'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Computational Molecular Design와 Agentic AI for Scientific Automation가 맞닿아, 'AI Scientists Fail Without Strong Implementation Capability'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Computational Molecular Design와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'Accelerating drug discovery with artificial: a whole-lab orchestration and scheduling system for self-driving labs'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구LLM 벤치마크 설계의 기초적 방법론을 제공하는 연구
기반 연구SPECTER2 유사도 0.92로 Computational Molecular Design와 Molecular Simulation and Generative Modeling가 맞닿아, 'Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근특정 도메인 특화 LLM 벤치마크 구축이라는 동일 문제를 다루는 대안적 연구로 판단된다.
다른 접근생물학적 특화 LLM 벤치마크를 다른 도메인으로 구성한 연구
후속 연구peptide 관련 LLM 평가를 확장한 벤치마크 연구
응용 사례LLM의 특정 도메인 이해력 평가를 실제 문제에 적용한 사례
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드