저자: Takayuki Shinohara | 날짜: 2026 | URL: https://openreview.net/forum?id=J263rcTa09 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Figure 1.
LLM 에이전트를 이용한 구조해석(structural analysis) 평가를 위해 OpenSeesPy 중심의 벤치마크 OpenSeesAgentBench를 제안하며, domain knowledge, code generation, end-to-end workflow 실행 역량을 분리 평가하는 contract-first 프레임워크와 reference 실행 스택을 함께 공개한다.
Figure 4. Qualitative execution examples from OpenSeesBuildingBench calibration profile, shown for cases 01, 04, 08, and
Figure 3. Deterministic dataset construction and split gen-
총평: 벤치마크 구성과 agent 역량 평가를 명확히 분리하려는 시도는 구조공학 AI 에이전트 평가 분야에 중요한 방법론적 기여를 하며, 공개된 reference 스택과 calibration 결과는 향후 비교 연구의 신뢰할 수 있는 기반이 될 것으로 보인다. 다만 현재는 단일 시스템의 calibration 검증에 그쳐 있어, 실제 다양한 agent 아키텍처 간 비교를 통한 실질적 유용성 입증은 후속 연구를 기다려야 한다.