An illustration of evaluation on InfiAgent-DABench. A LLM-based agent is prompted with a data analysis question and
InfiAgent-DABench는 LLM 기반 에이전트가 CSV 파일에 대해 end-to-end로 데이터 분석 작업을 수행하는 능력을 평가하기 위한 최초의 벤치마크로, 폐쇄형(closed-form) 질의응답 형식과 자동 채점 방식을 제안한다.
The workflow of DAEval construction. Data analysis questions are generated with GPT-4 based on the description of CSV
The workflow of DAEval construction. Data analysis questions are generated with GPT-4 based on the description of CSV
총평: LLM 에이전트의 데이터 분석 능력을 체계적이고 자동화된 방식으로 평가할 수 있는 최초의 벤치마크를 제시했다는 점에서 실용적 가치가 크며, DAAgent를 통해 벤치마크의 활용 가능성까지 보여준 완성도 높은 연구다.