We find that LLMs
본 연구는 STEM 분야의 design problem 창의성 평가에서 인간 전문가와 LLM이 originality를 판단할 때 어떤 인지적 과정과 세부 요인(uncommonness, remoteness, cleverness)에 의존하는지를 예시(example) 제공 여부에 따라 비교 분석한다.
CLAUDE-3.5-SONNET’s ratings generally agreed
Pearson correlations among pairwise Likert ratings
총평: 인간과 LLM의 창의성 평가 과정을 세밀한 facet 단위로 병렬 비교한 실증적이고 시의적절한 연구로, LLM을 창의성 평가자로 활용할 때 발생할 수 있는 잠재적 편향(예: facet 동질화)을 구체적으로 드러낸 점이 인상적이다. 다만 표본 규모와 도메인 일반화 측면에서 추가 검증이 필요하다.