AgentSLR: Automating Systematic Literature Reviews in Epidemiology with Agentic AI

저자: Shreyansh Padarha, Ryan Othniel Kearns, Tristan Myles Naidoo, Lingyi Yang, Łukasz Borchmann, Piotr Blaszczyk, Christian Morgenstern, Ruth McCabe, Sangeeta Bhatia, Philip Torr, Jakob Nicolaus Foerster, Scott A. Hale, Thomas Rawson, Anne Cori, Elizaveta Semenova, Adam Mahdi | 날짜: 2026 | URL: https://openreview.net/forum?id=qrR2mrP0TH 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. AgentSLR for assisting and evaluating large language models on systematic literature reviews in epidemiology (

AgentSLR은 article retrieval부터 screening, data extraction, report synthesis까지 systematic literature review(SLR)의 전체 워크플로우를 agentic LLM/LRM 파이프라인으로 자동화하는 open-source harness이며, 9개 WHO priority pathogen에 대한 역학 SLR에서 인간 전문가 수준의 성능을 유지하면서 리뷰 시간을 약 7주에서 20시간으로 58배 단축시킴을 보인다.

Motivation

Achievement

Figure 5

Figure 5. Model ablation results with AgentSLR across all pipeline stages. Averages are computed over the pathogens eval

  1. 전체 워크플로우 자동화: retrieval부터 report synthesis까지 SLR 전 과정을 커버하는 open-source agentic harness AgentSLR을 구축함.
  2. 대규모 실증 검증: 9개 WHO priority pathogen(예: Ebola, Lassa, SARS-CoV-1, Zika 등)에 대한 PERG의 expert-curated 라벨 대비 성능을 검증하여 인간 연구자와 comparable한 성능을 달성함.
  3. 획기적 속도 향상: 리뷰 소요 시간을 약 7주에서 20시간으로 58배 단축함.
  4. 모델 비교 분석: gpt-oss-120b, GPT-5.2, Kimi-K2.5, GLM-4.7, DeepSeek-V3.2 등 5개 frontier reasoning model을 파이프라인 각 단계별로 ablation하여, 성능이 모델 크기나 추론 비용보다 모델별 고유 역량(distinctive capabilities)에 좌우됨을 규명함.
  5. 실패 모드 규명: human-in-the-loop validation을 통해 주요 실패 모드를 식별함.

How

Figure 1

Figure 1. AgentSLR for assisting and evaluating large language models on systematic literature reviews in epidemiology (

Originality

Limitation & Further Study

Evaluation

Novelty: 4/5 Technical Soundness: 4/5 Significance: 5/5 Clarity: 4/5 Overall: 4/5

총평: SLR 전 과정을 아우르는 최초 수준의 통합 agentic 자동화 시스템을 고위험 실전 도메인(역학)에서 대규모로 검증했다는 점에서 실용적 임팩트가 크며, 모델 선택에 대한 실증적 통찰도 유용하다. 다만 open-access 데이터 편중과 일부 병원체 제외로 인한 검증 범위의 제약은 향후 보완이 필요하다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.95로 Biomedical AI Knowledge Systems와 AI-Driven Drug and Materials Discovery가 맞닿아, 'Bio-SIEVE: Exploring Instruction Tuning Large Language Models for Systematic Review Automation'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.94로 Biomedical AI Knowledge Systems와 AI-Assisted Academic Scholarly Communication가 맞닿아, 'Automated review generation method based on large language models'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구article retrieval 기반 리뷰 파이프라인의 기초적 구조를 공유한다.
기반 연구SPECTER2 유사도 0.95로 Biomedical AI Knowledge Systems와 LLM Benchmarking and Agent Evaluation가 맞닿아, 'AI-Researcher: Autonomous Scientific Innovation'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구유전체 변이 근거 추출이라는 실제 문제에 LLM을 적용한다.
다른 접근전체 연구 워크플로우를 agentic 파이프라인으로 자동화하는 유사한 접근이다.
다른 접근human-in-the-loop LLM 시스템을 통한 데이터 추출의 다른 구현 방식을 제시한다.
다른 접근두 논문 모두 역학(epidemiology) 분야에서 LLM 기반 agentic 파이프라인을 구축하지만, EpiAgent는 시뮬레이터 자동 합성을, AgentSLR은 문헌 검토 자동화를 목표로 한다.
응용 사례동일한 역학 도메인에서 agentic 자동화를 적용한 유사 연구이다.
다른 접근evidence sufficiency 판단 능력을 진단하는 유사한 벤치마크 접근이다.
다른 접근Clinical AI Lifecycle과 유사한 구현과학 프레임워크 매핑을 다루어 확장 관계로 볼 수 있음.
다른 접근온라인 데이터 기반 질병 예측의 유사한 사례 연구이다
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드