Cross-modal transfer learning for mapping bulk transcriptomes at cellular level

저자: Aarthi Venkat, Zhiwen Jiang, Daniel Marbach, Nir Hacohen, Marinka Zitnik | 날짜: 2026 | URL: https://openreview.net/forum?id=5bHHy81McV 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. (a) From large-scale single-cell RNA-seq data and sim-

POPPY는 ontology 기반 contrastive learning을 이용해 single-cell과 bulk transcriptomic foundation model을 정렬하고, cell-type-aware한 bulk patient embedding을 구성하는 프레임워크이다. 이를 통해 fine-tuning 없이도 bulk tumor transcriptome으로부터 세포 유형 및 gene program 구조를 복원한다.

Motivation

Achievement

Figure 3

Figure 3. (a) POPPY preserves cell type and gene program propor-

  1. Cell Ontology 기반 patient 유사도 정의: cell type composition을 Cell Ontology 그래프 상의 신호로 표현하고 Earth Mover's Distance(optimal transport)를 이용해 생물학적으로 구조화된 patient-patient 유사도 행렬 S를 정의하였다.
  2. Cross-modal contrastive alignment 프레임워크 구축: single-cell encoder와 bulk encoder를 multi-positive KL divergence 기반 대조 학습으로 정렬하여, fine-tuning 없이 bulk tumor transcriptome으로부터 세포 유형 및 gene program 구조를 복원하였다.
  3. Cross-cohort positive sampler 제안: cohort/assay 간 일반화를 향상시키기 위한 구조화된 positive sampling 전략을 도입하여 cohort 및 assay integration 성능을 개선하였다.
  4. 임상 적용 검증: 1,458개 single-cell tumor sample(3CA)과 1,286개 sorted bulk profile(Zaitsev et al., 2022)로 학습하여, 6개 bulk melanoma cohort(330개 종양)에서 immunotherapy response 예측 성능을 bulk-only encoder 및 contrastive 학습 없는 fine-tuned model보다 향상시켰고, T cell·NK cell·memory B cell을 responder profile과 연관된 세포로 식별하여 기존 문헌과 일치하는 결과를 도출하였다.

How

Figure 2

Figure 2. (a) POPPY improves fidelity of cross-modal alignment

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Single-cell과 bulk foundation model을 ontology 기반 optimal transport로 연결하는 독창적인 cross-modal contrastive learning 프레임워크로, 임상적으로 의미 있는 해석 가능한 bulk 임베딩을 제공한다는 점에서 임팩트가 크다. 다만 melanoma 중심의 검증에 국한되어 있어 다양한 질환 및 대규모 코호트로의 확장 검증이 향후 중요한 과제로 남는다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.93로 Multimodal Biomedical Data Fusion와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Integrated analysis of multimodal single-cell data'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 Multimodal Biomedical Data Fusion와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Efficient fine-tuning of single-cell foundation models enables zero-shot molecular perturbation prediction'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 Multimodal Biomedical Data Fusion와 LLMs for Molecular Biology & Chemistry가 맞닿아, 'CASSIA: a multi-agent large language model for reference free, interpretable, and automated cell annotation of single-cell RNA-sequencing data'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구유전자 발현 데이터 분석에 self-supervised 학습을 적용한 사례이다
기반 연구single-cell foundation model 학습 방법론이 POPPY의 cell-type-aware embedding 구성에 이론적 기반을 제공한다.
다른 접근단일세포 전사체 데이터의 세포 유형 주석 문제를 다루는 동일 분야 연구
기반 연구자연적 augmentation 개념을 다른 생물학적 데이터에 확장 적용한다.
기반 연구SPECTER2 유사도 0.94로 Multimodal Biomedical Data Fusion와 AI-Driven Drug and Materials Discovery가 맞닿아, 'AetherCell: A generative engine for virtual cell perturbation and in vivo drug discovery'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근bulk-to-single-cell 매핑을 위한 다른 contrastive learning 전략을 제안한다.
다른 접근single-cell과 bulk transcriptome 정렬을 위한 다른 contrastive learning 접근법을 제시함
다른 접근파운데이션 모델의 임베딩 품질을 층별로 평가하는 유사한 문제의식을 공유한다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드