Seed-coder: Let the code model curate data for itself

저자: ByteDance Seed, Yuyu Zhang, Jing Su, Yifan Sun, Chenguang Xi, Xia Xiao, Zheng Shen, A. Q. Zhang, Kaibo Liu, Daoguang Zan, Tao Sun, J. Zhu, Shijie Xin, Dong Huang, Y. Bai, Lixin Dong, C. J. Li, Jianchong Chen, Hao Zhou, Yifan Huang | 날짜: 2025 | URL: https://arxiv.org/abs/2506.03524 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

Essence

Figure 2

Figure 2. Processing pipeline for pretraining data. We collected data from GitHub and web archives.

Seed-Coder는 손수 작성한 필터링 규칙 대신 LLM 기반 점수 매기기 및 필터링을 사용하는 모델 중심 데이터 파이프라인으로 코드 사전학습 데이터를 자동으로 큐레이션하며, 8B 규모의 기본·명령·추론 모델을 제시한다.

Motivation

Achievement

Figure 1

Figure 1. Benchmark performance of instruct and reasoning variants of Seed-Coder-8B.

How

Figure 2

Figure 2. Processing pipeline for pretraining data. We collected data from GitHub and web archives.

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 3/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: Seed-Coder는 코드 사전학습 데이터 큐레이션의 패러다임을 hand-crafted 규칙에서 LLM 기반 자동화로 전환하며, 실용적이면서도 강력한 8B 모델을 통해 이 접근법의 효과를 명확히 입증한다. 확장 가능성과 객관성 측면에서 중요한 기여이나, 대규모 모델 및 필터 편향 분석에 대한 추가 탐구가 필요하다.

같이 보면 좋은 논문

다른 접근코드 모델 학습을 위한 데이터 큐레이션을 다른 방식으로 접근한다.
후속 연구대규모 토큰 기반 코드 모델 학습의 이론적 기초를 제공한다.
후속 연구동일한 코드 데이터 큐레이션 파이프라인을 다루는 매우 밀접한 후속 또는 확장 연구이다.
후속 연구SPECTER2 유사도 0.89로 Multimodal Biomedical Data Fusion와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Seed-coder: Let the code model curate data for itself'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
후속 연구SPECTER2 유사도 0.91로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Seed-coder: Let the code model curate data for itself'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
후속 연구SPECTER2 유사도 0.92로 Formal Proof Verification Automation와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Seed-coder: Let the code model curate data for itself'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
후속 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Seed-coder: Let the code model curate data for itself'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
후속 연구SPECTER2 유사도 0.93로 LLM Agent Reasoning Training와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Seed-coder: Let the code model curate data for itself'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
후속 연구SPECTER2 유사도 0.91로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Seed-coder: Let the code model curate data for itself'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
응용 사례LLM 기반 데이터 큐레이션 기법을 코드 모델 학습에 적용한다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드