Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment

저자: Yanjun Shao, Yundi Chen, Yashvi Patel, Aurelien Pelissier, María Rodríguez Martínez | 날짜: 2026 | URL: https://openreview.net/forum?id=oHaY9DxcTN 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Overview of LOGICA: pretrained biological language models are coupled by cross-modal adapters that preserve na

사전학습된 biological language model(BioFM)의 native token head를 유지하면서, output logit 공간에서 직접 contrastive learning을 수행하여 context-conditioned 확률로 변이(variant) 및 결합(binding) 랭킹을 수행하는 프레임워크 LogiCA를 제안한다.

Motivation

Achievement

Figure 2

Figure 2. Two scaling regimes for protein–ligand LOGICA. (A) Held-out likelihood-margin trajectories during pretraining

  1. Logit-space contrastive alignment 프레임워크 제안: pooled embedding이나 별도 예측 head 없이, native token head가 산출하는 context-conditioned log-likelihood를 직접 compatibility score로 사용하는 LogiCA를 제안하였다.
  2. Mutation-local variant ranking에 대한 이론적 정당화: 동일 mutation site를 공유하는 변이 쌍에 대해 wild-type 공통 항이 상쇄되어 Bradley–Terry 형태의 랭킹 확률로 환원됨을 명제(Proposition 2.1)로 증명하였다.
  3. 다중 도메인 실증 성능 향상: protein–ligand binding, TCR–peptide activity, drug-conditioned resistance prediction 세 과제에서 matched latent-contrastive 및 conditional-MLM baseline 대비 상당한 성능 향상을 보였으며, held-out-gene 단일 변이 약물 저항성 예측에서 AUC를 latent-space baseline의 ~0.55에서 ~0.65로 개선하였다.
  4. 서로 다른 vocabulary/tokenizer 간 정렬: 공유 tokenizer, decoder, embedding space 없이도 서로 다른 vocabulary를 가진 모델(ESM-2, TCRLang, SELFormer) 간 sparse paired data 정렬을 가능하게 하였다.

How

Figure 1

Figure 1. Overview of LOGICA: pretrained biological language models are coupled by cross-modal adapters that preserve na

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: 사전학습된 biological language model의 핵심 자산인 token-level likelihood interface를 훼손하지 않으면서 context-conditioned contrastive learning을 가능케 한 개념적으로 신선하고 실용적인 기여이며, 다양한 biological modality에 걸친 실증 결과가 이를 뒷받침한다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.92로 Multimodal Biomedical Data Fusion와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'Effective gene expression prediction from sequence by integrating long-range interactions'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구생물학적 언어모델의 contrastive learning 기반을 제공한다.
기반 연구SPECTER2 유사도 0.93로 Multimodal Biomedical Data Fusion와 AI-Driven Drug and Materials Discovery가 맞닿아, 'CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Multimodal Biomedical Data Fusion와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'Partially shared multi-modal embedding learns holistic representation of cell state (APOLLO)'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근세포 상태 표현 학습의 다른 통합 방식을 다룬다.
기반 연구SPECTER2 유사도 0.93로 Multimodal Biomedical Data Fusion와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Linear-time prediction of proteome-scale microbial protein interactions'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Multimodal Biomedical Data Fusion와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Cross-Attention Over RNA And Protein Sequences Enables Generalizable Interaction Prediction'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
후속 연구output logit 공간에서의 대조학습을 확장한 연구이다.
응용 사례생물학적 시퀀스 모델을 실제 변이 분석에 적용한다.
응용 사례사전학습된 BioFM을 다중 modality 문제에 적용한다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드