저자: Akash Ghosh, Srivarshinee Sridhar, Raghav Kaushik Ravi, Muhsin Muhsin, Sriparna Saha, Chirag Agarwal | 날짜: 2026 | URL: https://openreview.net/forum?id=bMsBArF8MT 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Fig. 1). We employ a novel two-step approach to generate
CLINIC은 의료 분야 language model의 신뢰성(trustworthiness)을 truthfulness, fairness, safety, robustness, privacy 5개 축에서 15개 언어, 18개 task, 28,800개 샘플로 평가하는 최초의 종합적 다국어 benchmark이다. 13개 모델(소형/대형 open-weight, 의료 특화, reasoning 모델 포함)을 평가한 결과, 사실 정확성 부족, 인구통계·언어 그룹 간 편향, privacy breach 및 adversarial attack에 대한 취약성이 공통적으로 드러났다.
Figure 3. Average (across false confidence, false question, and none of the above test) model hallucination accuracy (↑)
Figure 2. Construction of CLINIC. Step 1 involves data collection and mapping English samples to their corresponding mul
총평: 의료 AI의 글로벌 배포에 있어 반드시 필요한 다국어 trustworthiness 평가의 공백을 체계적으로 메운 선구적이고 실용적 가치가 매우 높은 benchmark 연구이다. 데이터 구축의 엄밀성과 전문가 검증 과정이 돋보이며, 향후 다국어 의료 AI 안전성 연구의 표준 참조점이 될 잠재력이 크다.