BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

저자: Haitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun, Qihao Zheng, Mianxin Liu, Chi Zhang, Wanli Ouyang, Chunfeng Song, Changqing Zhang, Jiamin Wu | 날짜: 2026 | URL: https://openreview.net/forum?id=nJxailqsUW 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 2

Figure 2. An overview of BrainJanus. The input data, regardless of its modality, is tokenized into a shared token space

뇌 활동, 시각, 언어를 하나의 discrete token space로 통합하여 brain encoding(image/text→brain)과 decoding(brain→image/text)을 단일 autoregressive 모델로 수행하는 최초의 통합 프레임워크 BrainJanus를 제안한다.

Motivation

Achievement

Figure 3

Figure 3. Qualitative comparison of brain caption decoding results. GroundTruth image captions are compared with caption

  1. 최초의 통합 brain-vision-language 모델: brain encoding과 decoding, 총 4가지 task(brain-to-image, brain-to-text, image-to-brain, text-to-brain)를 모두 지원하는 최초의 unified model을 제안했다(Table 1 비교에서 기존 방법들은 이 중 일부만 지원).
  2. 다양한 벤치마크에서 우수한 성능: brain decoding 및 encoding 벤치마크 전반에서 task-specific 모델들을 능가하는 competitive 성능을 달성했다.
  3. Zero-shot 일반화: 통합 multi-task 학습을 통해 cross-modal 지식 전이가 촉진되어, task-agnostic한 강건한 representation으로 zero-shot generalization 능력을 확인했다.
  4. 생물학적 해석 가능성 보존: 생성된 fMRI 신호가 해석 가능한 cortical topography와 biological variability를 보존함을 보여, 모델이 의미 있는 neural representation을 학습했음을 입증했다.

How

Figure 2

Figure 2. An overview of BrainJanus. The input data, regardless of its modality, is tokenized into a shared token space

Originality

Limitation & Further Study

Evaluation

Novelty: 5/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: 뇌를 omni-modal 시스템으로 보는 신경과학적 관점을 discrete tokenization과 unified autoregressive 구조로 구현한 참신하고 야심찬 시도로, brain encoding/decoding 통합이라는 오랜 문제에 새로운 패러다임을 제시한다는 점에서 의의가 크다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.92로 Multimodal Biomedical Data Fusion와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'GPT-4 Technical Report'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.91 기준으로 'BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language'의 AI4S 방법론을 'Deepseek-v3 technical report'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구동일한 저자 계열의 통합 brain encoding/decoding 모델로 확장된 연구
후속 연구ECoG 기반 모터 디코딩을 통합 모델의 세부 응용으로 확장
기반 연구diffusion transformer를 뇌 신호 생성에 적용한 유사 연구이다.
다른 접근멀티모달 통합 생성 모델로서 유사한 문제를 다룸
기반 연구지식 기반 의미 프로파일링 모듈을 확장하여 적용한 후속 연구이다.
기반 연구SPECTER2 유사도 0.92로 Multimodal Biomedical Data Fusion와 AI-Driven Drug and Materials Discovery가 맞닿아, 'CLM-X: A multimodal single-cell foundation model with flexible multi-way Transformer for unified scRNA-seq and scATAC-seq analysis'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.91 기준으로 'BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language'의 AI4S 방법론을 'Generative AI and the Foundation Model Era: A Comprehensive Review'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구SPECTER2 유사도 0.92로 Multimodal Biomedical Data Fusion와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Generative machine learning unlocks the first proteome-wide image of human cells'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근뇌-이미지-언어 통합 표현을 다루는 대안적 모델
다른 접근brain-vision-language 통합 문제에 대한 다른 방법론 제시
다른 접근long-range dependency 포착을 위한 유사한 시퀀스 모델링 기법이다.
다른 접근뇌 신호 처리를 위한 다른 아키텍처 접근법
후속 연구계층적 시각 처리 기반 신경 인코딩의 이론적 토대 제공
응용 사례통합 뇌 인코딩/디코딩 모델의 실제 응용 사례
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드