저자: Sihui Dai, Mann Patel | 날짜: 2026 | URL: https://openreview.net/forum?id=aS7CLe6qAt 📄 PDF
라이선스: OpenReview 공개(오픈액세스)
Figure 1. Benign vs. Harmful Compliance Demonstration. In a
본 논문은 in-context demonstration을 통한 jailbreaking에서, benign compliance demonstration(무해한 요청+도움이 되는 응답)과 harmful compliance demonstration(유해한 요청+거절하지 않는 응답)을 섞은 mixed-demonstration context가 모델의 harmful query에 대한 compliance에 어떻게 영향을 미치는지 세 가지 가설(total-count, harmful-count, joint-count hypothesis)로 검증한다. 네 개의 모델을 대상으로 실험하여 benign demonstration이 모델에 따라 harmful compliance를 줄이거나(dilution) 늘리는(amplification) 상반된 효과를 낼 수 있음을 보이고, 이 차이가 preference optimization(DPO) 학습 단계 여부에 기인함을 밝힌다.
Figure 4. Impact of varying benign compliance demonstra-
Figure 4. Impact of varying benign compliance demonstra-
총평: demonstration-based jailbreaking의 "작동 여부"를 넘어 "작동 방식"을 규명하려는 시도로서 방법론적으로 명확하고 실용적 함의가 큰 연구이나, 모델 수와 학습 단계 비교 범위가 제한적이어서 일반화 가능성에 대한 추가 검증이 필요하다.