Essence
Figure 1. CI width versus the number of evaluated samples. Each trajectory is averaged over 50 seeds and surrogate–targe
CELEUS는 e-process 기반의 anytime-valid confidence interval(CI)을 구성하여 LLM 평가에 통계적으로 엄밀한 인증(certification)을 제공하면서, uncertainty-guided sampling과 surrogate-assisted approximation을 결합해 더 적은 평가 샘플로 목표 정밀도에 도달하는 프레임워크이다.
Evaluation
Novelty: 4/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5
총평: e-process를 활용해 LLM 평가의 anytime-validity 문제를 이론적으로 엄밀하게 해결하면서도 실용적 효율성까지 확보한 견고한 연구로, certifiable evaluation 분야에 의미 있는 기여를 한다.
같이 보면 좋은 논문
기반 연구SPECTER2 유사도 0.93로 LLM Reasoning and Safety Benchmarks와 Formal Methods and Computational Reasoning가 맞닿아, 'TrustLLM: Trustworthiness in Large Language Models'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근LLM 신뢰성 평가를 위한 다른 벤치마크 프레임워크
기반 연구SPECTER2 유사도 0.92로 LLM Reasoning and Safety Benchmarks와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'Mind the gap: Examining the self-improvement capabilities of large language models'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.92로 LLM Reasoning and Safety Benchmarks와 Scientific AI for Physics and Environment가 맞닿아, 'RBF++: Quantifying and optimizing reasoning boundaries across measurable and unmeasurable capabilities for chain-of-thought reasoning'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구anytime-valid inference 및 e-process 이론이라는 통계적 기반을 공유
기반 연구anytime-valid 신뢰구간의 이론적 기반을 제공한다.
기반 연구비용을 고려한 최적화 정책을 확장하여 적용한 연구이다.
기반 연구Ising-type energy minimization을 다른 형태의 신호 통합에 확장 적용한 연구이다.
다른 접근동료평가에서의 LLM 오남용 탐지를 위한 대안적 방법을 제시한다.
다른 접근uncertainty-guided sampling 기반 LLM 평가라는 유사한 접근 방식
다른 접근sequential testing 문제를 다른 방식으로 접근한 대안적 프레임워크
다른 접근LLM 평가를 위한 통계적으로 엄밀한 신뢰구간 구성이라는 유사한 목표를 다른 방법으로 접근
후속 연구KOH 다중 신뢰도 GP 모델의 이론적 기반을 공유한다.
후속 연구functional score test의 통계적 기반을 제공하는 선행 연구이다.
후속 연구surface-level과 approach-level diversity 구분의 이론적 토대