저자: Aarav Sinha | 날짜: 2026 | URL: https://openreview.net/forum?id=yyEEwEAa5a 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Figure 1. Digital-fly discovery loop. The agent observes compressed neural and behavioral readouts, chooses budgeted
이 논문은 whole-brain Drosophila 모델(digital fly)을 이용해 AI agent가 신경 메커니즘을 실제로 발견할 수 있는지 평가하는 benchmark를 제안하고, feeding/grooming circuit 관련 6개 pilot task로 이를 검증한다. planning agent가 memory와 budgeted experimentation을 결합했을 때 one-shot 및 no-memory/memory-only variant보다 Node F1에서 우수하며, named-versus-anonymized 조건에서 shortcut robustness 차이를 드러낸다.
Figure 1. Digital-fly discovery loop. The agent observes compressed neural and behavioral readouts, chooses budgeted
Figure 1. Digital-fly discovery loop. The agent observes compressed neural and behavioral readouts, chooses budgeted
총평: AI agent의 과학적 discovery 능력을 mechanistic reasoning 관점에서 평가하는 신선한 benchmark 설계를 제시했으나, pilot 규모가 작고 일부 task family(edge-level attribution 등)가 미완성이라 정식 leaderboard보다는 초기 proof-of-concept 성격이 강하다.