VisualGPTScore: Visio-Linguistic Reasoning with Multimodal Generative Pre-Training Scores

Lin, Zhiqiu; Chen, Xinyue; Pathak, Deepak; Zhang, Pengchuan; Ramanan, Deva

Full-text links:

Download:

Current browse context:

cs.CV

< prev | next >

new | recent | 2306

Computer Science > Computer Vision and Pattern Recognition

Title: VisualGPTScore: Visio-Linguistic Reasoning with Multimodal Generative Pre-Training Scores

Authors: Zhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang, Deva Ramanan

(Submitted on 2 Jun 2023 (this version), latest version 1 Feb 2024 (v3))

Abstract: Vision-language models (VLMs) discriminatively pre-trained with contrastive image-text matching losses such as $P(\text{match}|\text{text}, \text{image})$ have been criticized for lacking compositional understanding. This means they might output similar scores even if the original caption is rearranged into a different semantic statement. To address this, we propose to use the ${\bf V}$isual ${\bf G}$enerative ${\bf P}$re-${\bf T}$raining Score (${\bf VisualGPTScore}$) of $P(\text{text}|\text{image})$, a $\textit{multimodal generative}$ score that captures the likelihood of a text caption conditioned on an image using an image-conditioned language model. Contrary to the belief that VLMs are mere bag-of-words models, our off-the-shelf VisualGPTScore demonstrates top-tier performance on recently proposed image-text retrieval benchmarks like ARO and Crepe that assess compositional reasoning. Furthermore, we factorize VisualGPTScore into a product of the $\textit{marginal}$ P(text) and the $\textit{Pointwise Mutual Information}$ (PMI). This helps to (a) diagnose datasets with strong language bias, and (b) debias results on other benchmarks like Winoground using an information-theoretic framework. VisualGPTScore provides valuable insights and serves as a strong baseline for future evaluation of visio-linguistic compositionality.

Comments:	Website: this https URL Code: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2306.01879 [cs.CV]
	(or arXiv:2306.01879v1 [cs.CV] for this version)

Submission history

From: Zhiqiu Lin [view email]
[v1] Fri, 2 Jun 2023 19:19:43 GMT (8563kb,D)
[v2] Thu, 5 Oct 2023 04:12:28 GMT (13285kb,D)
[v3] Thu, 1 Feb 2024 18:22:25 GMT (10206kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2306.01879v1

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: VisualGPTScore: Visio-Linguistic Reasoning with Multimodal Generative Pre-Training Scores

Submission history