MAGVLT: Masked Generative Vision-and-Language Transformer

Kim, Sungwoong; Jo, Daejin; Lee, Donghoon; Kim, Jongmin

Full-text links:

Download:

Current browse context:

cs.CV

< prev | next >

new | recent | 2303

Computer Science > Computer Vision and Pattern Recognition

Title: MAGVLT: Masked Generative Vision-and-Language Transformer

Authors: Sungwoong Kim, Daejin Jo, Donghoon Lee, Jongmin Kim

(Submitted on 21 Mar 2023)

Abstract: While generative modeling on multimodal image-text data has been actively developed with large-scale paired datasets, there have been limited attempts to generate both image and text data by a single model rather than a generation of one fixed modality conditioned on the other modality. In this paper, we explore a unified generative vision-and-language (VL) model that can produce both images and text sequences. Especially, we propose a generative VL transformer based on the non-autoregressive mask prediction, named MAGVLT, and compare it with an autoregressive generative VL transformer (ARGVLT). In comparison to ARGVLT, the proposed MAGVLT enables bidirectional context encoding, fast decoding by parallel token predictions in an iterative refinement, and extended editing capabilities such as image and text infilling. For rigorous training of our MAGVLT with image-text pairs from scratch, we combine the image-to-text, text-to-image, and joint image-and-text mask prediction tasks. Moreover, we devise two additional tasks based on the step-unrolled mask prediction and the selective prediction on the mixture of two image-text pairs. Experimental results on various downstream generation tasks of VL benchmarks show that our MAGVLT outperforms ARGVLT by a large margin even with significant inference speedup. Particularly, MAGVLT achieves competitive results on both zero-shot image-to-text and text-to-image generation tasks from MS-COCO by one moderate-sized model (fewer than 500M parameters) even without the use of monomodal data and networks.

Comments:	CVPR 2023
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2303.12208 [cs.CV]
	(or arXiv:2303.12208v1 [cs.CV] for this version)

Submission history

From: Sungwoong Kim [view email]
[v1] Tue, 21 Mar 2023 21:49:39 GMT (39376kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2303.12208

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: MAGVLT: Masked Generative Vision-and-Language Transformer

Submission history