저자: Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas Beyer | 날짜: 2023-03-27 | URL: https://arxiv.org/abs/2303.15343 📄 PDF
라이선스: CC BY
Figure 1: Efficient loss implementation demonstrated via a mock setup with 3 devices and a global batch size of 12. There
Language-Image Pre-training을 위해 softmax 정규화 대신 pairwise sigmoid loss를 제안하며, 이는 배치 크기와 무관하게 작동하여 메모리 효율성을 개선하고 작은 배치 크기에서 더 나은 성능을 달성한다.
Figure 2: The effect of pre-training batch size. Left: SigLiT results, trained for 18B seen examples. Sigmoid loss outpe
Figure 1: Efficient loss implementation demonstrated via a mock setup with 3 devices and a global batch size of 12. There
총평: Sigmoid loss를 통해 language-image pre-training의 효율성과 확장성을 동시에 개선한 우수한 연구로, 실무적 접근 가능성을 크게 높이며 배치 크기의 영향에 대한 중요한 통찰을 제공한다.