Exploring Semantic Relationships for Unpaired Image Captioning

Liu, Fenglin; Gao, Meng; Zhang, Tianhao; Zou, Yuexian

Full-text links:

Download:

Current browse context:

cs.CV

< prev | next >

new | recent | 2106

Change to browse by:

Computer Science > Computer Vision and Pattern Recognition

Title: Exploring Semantic Relationships for Unpaired Image Captioning

Authors: Fenglin Liu, Meng Gao, Tianhao Zhang, Yuexian Zou

(Submitted on 20 Jun 2021 (v1), last revised 17 Aug 2021 (this version, v2))

Abstract: Recently, image captioning has aroused great interest in both academic and industrial worlds. Most existing systems are built upon large-scale datasets consisting of image-sentence pairs, which, however, are time-consuming to construct. In addition, even for the most advanced image captioning systems, it is still difficult to realize deep image understanding. In this work, we achieve unpaired image captioning by bridging the vision and the language domains with high-level semantic information. The motivation stems from the fact that the semantic concepts with the same modality can be extracted from both images and descriptions. To further improve the quality of captions generated by the model, we propose the Semantic Relationship Explorer, which explores the relationships between semantic concepts for better understanding of the image. Extensive experiments on MSCOCO dataset show that we can generate desirable captions without paired datasets. Furthermore, the proposed approach boosts five strong baselines under the paired setting, where the most significant improvement in CIDEr score reaches 8%, demonstrating that it is effective and generalizes well to a wide range of models.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2106.10658 [cs.CV]
	(or arXiv:2106.10658v2 [cs.CV] for this version)

Submission history

From: Fenglin Liu [view email]
[v1] Sun, 20 Jun 2021 09:10:11 GMT (2042kb,D)
[v2] Tue, 17 Aug 2021 15:29:46 GMT (2044kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2106.10658

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: Exploring Semantic Relationships for Unpaired Image Captioning

Submission history