Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
Jinhao Li, Haopeng Li, Sarah Erfani, Lei Feng, James Bailey, and Feng Liu. (2024). "Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models." Proceedings of the International Conference on Machine Learning (ICML).