Multimodal Chain-of-Thought Reasoning in Language Models

Zhang, Zhuosheng; Zhang, Aston; Li, Mu; Zhao, Hai; Karypis, George; Smola, Alex

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2302

Computer Science > Computation and Language

Title: Multimodal Chain-of-Thought Reasoning in Language Models

Authors: Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, Alex Smola

(Submitted on 2 Feb 2023 (v1), revised 9 Feb 2023 (this version, v2), latest version 17 Feb 2023 (v4))

Abstract: Large language models (LLMs) have shown impressive performance on complex reasoning by leveraging chain-of-thought (CoT) prompting to generate intermediate reasoning chains as the rationale to infer the answer. However, existing CoT studies are mostly isolated in the language modality with LLMs, where LLMs are hard to deploy. To elicit CoT reasoning in multimodality, a possible solution is to fine-tune small language models by fusing the vision and language features to perform CoT reasoning. The key challenge is that those language models tend to generate hallucinated reasoning chains that mislead the answer inference. To mitigate the effect of such mistakes, we propose Multimodal-CoT that incorporates vision features. The framework separates the rationale generation and answer inference into two stages. By incorporating the vision features in both stages, the model is able to generate effective rationales that contribute to answer inference. With Multimodal-CoT, our model under 1 billion parameters outperforms the previous state-of-the-art LLM (GPT-3.5) by 16% (75.17%->91.68%) on the ScienceQA benchmark and even surpasses human performance. Code is publicly available at this https URL

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2302.00923 [cs.CL]
	(or arXiv:2302.00923v2 [cs.CL] for this version)

Submission history

From: Aston Zhang [view email]
[v1] Thu, 2 Feb 2023 07:51:19 GMT (421kb,D)
[v2] Thu, 9 Feb 2023 02:10:36 GMT (421kb,D)
[v3] Wed, 15 Feb 2023 19:20:15 GMT (416kb,D)
[v4] Fri, 17 Feb 2023 04:35:55 GMT (477kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2302.00923v2

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Multimodal Chain-of-Thought Reasoning in Language Models

Submission history