Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision

저자: Wei Du, Shubham Toshniwal, Branislav Kisacanin, Sadegh Mahdavi, Ivan Moshkov, George Armstrong, Stephen Ge, Edgar Minasyan, Feng Chen, Igor Gitman | 날짜: 2026 | URL: https://openreview.net/forum?id=FK9LEhmjof 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Scaling with model size and architecture on Nemotron-Math. Each panel reports pass@1 as a function of training

gpt-oss-120b의 multi-mode(high/medium/low) 및 Python TIR 유무 조합을 활용해 7.5M개의 장문 수학 추론 trace로 구성된 대규모 데이터셋 Nemotron-Math를 구축하고, 이를 통해 효율적인 long-context(최대 128K) distillation 학습 기법까지 함께 제시한 연구이다.

Motivation

Achievement

Figure 1

Figure 1. Scaling with model size and architecture on Nemotron-Math. Each panel reports pass@1 as a function of training

  1. 대규모 multi-mode 수학 추론 데이터셋 구축: gpt-oss-120b를 이용해 high/medium/low reasoning mode와 Python TIR 유무를 조합, 85K AoPS 문제와 262K StackExchange-Math 문제로부터 7.5M개의 최대 128K 토큰 길이 solution trace를 생성했다.
  2. 품질 검증 및 일반화 성능 개선: 통제된 비교실험을 통해 Nemotron-Math가 matched AoPS 문제에서 기존 OpenMathReasoning을 일관되게 능가하며, StackExchange-Math 추가가 HLE-Math 등에서 robustness와 generalization을 크게 향상시키면서도 competition benchmark 정확도를 유지함을 보였다.
  3. 효율적인 long-context 학습 전략 제안: sequential bucketed training 전략으로 128K context-length fine-tuning을 full-length 학습 대비 2–3배 가속하면서 정확도 손실을 1–3% 이내로 최소화했다.
  4. 모델 규모·아키텍처에 걸친 확장성 검증: Qwen3-8B와 Qwen3-30B-A3B에서 실험을 수행하여 두 모델 모두 유사한 최종 성능에 수렴함을 확인했고, high reasoning mode + Python TIR 설정에서 AIME 2024/2025에 대해 두 모델 모두 100% maj@16 정확도를 달성했다.

How

Figure 1

Figure 1. Scaling with model size and architecture on Nemotron-Math. Each panel reports pass@1 as a function of training

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Multi-mode teacher generation과 diverse problem source를 결합한 대규모 고품질 수학 추론 데이터셋 구축과 효율적인 long-context distillation 기법을 함께 제시한 실용적이고 완성도 높은 연구로, AIME 등에서의 SOTA 결과가 이를 뒷받침한다.

같이 보면 좋은 논문

기반 연구long-context distillation 기법의 이론적 기반을 제공한다.
기반 연구SPECTER2 유사도 0.92 기준으로 'Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision'의 AI4S 방법론을 'SciCode: A Research Coding Benchmark Curated by Scientists'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구SPECTER2 유사도 0.92로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'Advancing Mathematics Research with AI-Driven Formal Proof Search'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'Accelerating Scientific Research with Gemini: Case Studies and Common Techniques'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근장문 수학 추론 distillation 데이터셋 구축이라는 동일 목표를 다른 방식으로 달성하려는 연구로 보인다.
다른 접근장문 수학 추론 데이터셋 구축을 다른 모델 조합으로 접근한다.
다른 접근수학 정리 증명 벤치마크의 다른 도메인 특화 버전
다른 접근대규모 수학 추론 데이터셋 구축에 다른 distillation 전략을 사용한다.
후속 연구난이도 갭 문제 해결을 위한 커리큘럼 학습의 토대를 제공한다.
응용 사례동일한 수학 추론 trace 생성 방법을 실제 모델 학습에 적용한다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드