ScoreStop: Gradient-based early stopping using functional score tests

저자: Oliver J. Hines, Christian L. Hines | 날짜: 2026 | URL: https://openreview.net/forum?id=zVIk8t8OZO 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Single regression trajectory at η = 0.05. Top: validation

ScoreStop은 gradient boosting에서 각 반복마다 현재 예측기가 population risk minimizer인지 검정하는 functional score test 기반의 early stopping 규칙으로, scale-invariant 통계량이 귀무가설 하에서 점근적으로 chi-squared 분포를 따르도록 설계되었다.

Motivation

Achievement

Figure 3

Figure 3. Real-data benchmarks: excess test loss over the retrospective test-loss oracle. Markers are fold medians and v

  1. ScoreStop 통계량 제안: validation 데이터에서 계산되는 score function s(W;h,f)의 empirical mean과 variance를 이용해 Tn(h,f*) = n·En[s]^2/En[s^2] 형태의 scale-invariant 검정통계량을 구성하고, 이것이 귀무가설 하에서 χ²₁ 분포로 수렴함을 증명(Proposition 4.6).
  2. 대립가설 하 발산 증명: ⟨h, h(f)⟩≠0인 한 Tn이 대립가설 하에서 무한대로 발산함을 증명(Proposition 4.7), 방향으로 h=h(f)를 선택하는 것이 near-optimal임을 보임(Corollary 4.8).
  3. 일반화된 적용범위: smooth loss(squared-error, logistic, Poisson, multiclass softmax)뿐 아니라 LambdaRank 같은 implicit loss, Cox regression 같은 data-dependent loss(influence function 활용)에도 동일한 directional-derivative 원리로 확장 가능함을 보임.
  4. 경험적 검증: synthetic 실험과 regression, classification, count, ranking, quantile, survival 등 다양한 실제 tabular benchmark에서 ScoreStop이 기존 loss-based 방법과 경쟁력 있는 성능을 보임을 입증.

How

Figure 4

Figure 4. QQ plots of the score statistics in a ±5 iteration window around the median test-loss argmin iteration, for ea

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: ScoreStop은 gradient boosting의 early stopping을 통계적 가설 검정으로 원칙 있게 재구성한 독창적이고 이론적으로 탄탄한 연구로, implicit/data-dependent loss까지 아우르는 일반성이 돋보이나 유한 표본 상황과 다중 방향 검정으로의 확장에 대한 추가 검증이 필요하다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.90로 Statistical Causal Inference Methods와 Molecular Simulation and Generative Modeling가 맞닿아, 'Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.90로 Statistical Causal Inference Methods와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Mind the gap: Examining the self-improvement capabilities of large language models'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구functional score test의 통계적 기반을 제공하는 선행 연구이다.
기반 연구Bayesian optimization의 실제 응용에서 비용 효율성을 검증한 사례이다.
다른 접근gradient boosting의 조기 종료 문제를 다른 통계적 검정 방식으로 접근한다.
기반 연구SPECTER2 유사도 0.90로 Statistical Causal Inference Methods와 Molecular Simulation and Generative Modeling가 맞닿아, 'SamplingDesign: RNA design via continuous optimization with coupled variables and Monte-Carlo sampling'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.90로 Statistical Causal Inference Methods와 Applied Bibliometrics Across Domains가 맞닿아, 'Transitioning from invasive to liquid biopsy techniques: a bibliometric analysis and prospective insights on biomarkers in lupus nephritis'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
응용 사례early stopping 규칙을 실제 boosting 알고리즘에 적용한 사례이다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드