저자: Wei Liu, Peijie Yu, Michele Orini, Yali Du, Yulan He | 날짜: 2026 | URL: https://openreview.net/forum?id=NooBkSo7oz 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Figure 1. Left: Compared with previous tasks, DDR maximises exploration openness and agency, focusing on the direct eval
본 논문은 사전에 정해진 질문 없이 구조화된 데이터베이스로부터 스스로 탐색 목표를 설정하고 통찰을 추출하는 능력, 즉 investigatory intelligence를 평가하는 Deep Data Research(DDR) 과제와 이를 검증 가능하게 측정하는 대규모 벤치마크 DDR-Bench를 제안한다.
Figure 4. Inference-time scaling performance in DDR-Bench across different dimensions. The y-axis reports checklist accu
Figure 2. A case of Claude Sonnet 4.5’s trajectory and evaluation checklist in the MIMIC scenario of DDR-Bench. Verified
총평: Agentic LLM 평가에서 간과되어 온 investigatory intelligence라는 개념을 명확히 정의하고 이를 검증 가능하게 측정하는 벤치마크를 제시한 점에서 시의적절하고 독창적인 기여이며, 다만 체크리스트 방식의 근본적 한계와 도메인 일반화 검증이 향후 과제로 남는다.