Beyond Direct Gains: Matched Controls for Evaluating Concept Scaffolds

저자: Jiangshan He, Song Dai, Xiaolong Qiao, Jungang Li, Yibo Yan, Xuming Hu | 날짜: 2026 | URL: https://openreview.net/forum?id=R3ltJC0uZN 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Matched-control protocol: for each base problem, we construct five concept conditions, compute diagnostic cont

본 논문은 concept scaffold(개념 힌트)가 수학 문제 풀이 정확도를 높인다는 기존 평가 방식이 실제로는 semantic content, prompt format, plausible-but-wrong auxiliary information에 대한 민감도가 뒤섞인 under-identified 비교임을 지적하고, 이를 분리하기 위한 matched-control protocol을 제안한다.

Motivation

Achievement

Figure 1

Figure 1. Matched-control protocol: for each base problem, we construct five concept conditions, compute diagnostic cont

  1. Matched-control protocol 제안: N, F, D, M, W 다섯 조건과 G_format, G_helpful, G_mislead, G_corrupt 네 가지 diagnostic delta를 정의하여 concept scaffold의 효과를 semantic content, format, misleading sensitivity로 분해했다.
  2. 대규모 실증 평가: UGMathBench 기반 2,000개 base problem, 문제당 3개 randomized version, 조건당 6,000개 solver prompt로 총 9개 instruction-compatible model setting(Qwen2.5, Qwen2.5-Math, DeepSeek-R1 distilled, Qwen3, Gemma 3, Phi-4-mini, Llama 3.2)에 대해 평가했다.
  3. stable accuracy 개념 도입: 문제의 3개 randomized version 모두 정답이어야 correct로 인정하는 stable accuracy S를 주요 지표로 채택하여, 단일 표본 버전에서의 우연한 개선과 근본적 reasoning objective에 대한 안정적 개선을 구분했다.
  4. 핵심 실증 결과: 대부분의 setting에서 helpful concept이 format-only 대비 stable accuracy를 개선하지만 평균 gain은 미미함(G_helpful=+1.22점)을 보였고, 이를 semantic effect의 순수 추정치가 아닌 진단적 helpful-over-format contrast로 해석해야 함을 논증했다.
  5. Pseudo-concept 구축 파이프라인: UGMathBench에 CHAMP-style concept label이 없다는 한계를 극복하기 위해 Qwen2.5-72B-Instruct를 teacher로 활용한 concept 생성 및 reference-answer 기반 검증, 필터링 절차를 마련했다.

How

Figure 1

Figure 1. Matched-control protocol: for each base problem, we construct five concept conditions, compute diagnostic cont

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: reasoning scaffold 평가에서 흔히 간과되는 identification 문제를 명확히 제기하고 실용적인 matched-control protocol과 stable accuracy 지표를 제안한 견실한 연구로, workshop paper 수준에서 방법론적 기여가 뚜렷하지만 단일 도메인·단일 teacher model에 국한된 실증 범위는 향후 확장이 필요하다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.94로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'TheoremQA: A Theorem-driven Question Answering Dataset'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.94로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Curie: Toward rigorous and automated scientific experimentation with ai agents'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SHAPE 프레임워크의 semantic space 분석을 확장한 연구로 보인다.
기반 연구aggregation 기반 test-time 방법론을 다양한 reasoning 도메인으로 확장하는 연구이다.
기반 연구수학 추론 평가 벤치마크의 기초를 제공
다른 접근수학 추론 평가의 민감도 문제를 다른 방식으로 검증함
다른 접근concept scaffold 평가의 신뢰성 문제에 대한 다른 검증 방식을 제시한다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드