Consistency Analysis of ChatGPT

Jang, Myeongjun Erik; Lukasiewicz, Thomas

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2303

Computer Science > Computation and Language

Title: Consistency Analysis of ChatGPT

Authors: Myeongjun Erik Jang, Thomas Lukasiewicz

(Submitted on 11 Mar 2023 (v1), last revised 14 Nov 2023 (this version, v3))

Abstract: ChatGPT has gained a huge popularity since its introduction. Its positive aspects have been reported through many media platforms, and some analyses even showed that ChatGPT achieved a decent grade in professional exams, adding extra support to the claim that AI can now assist and even replace humans in industrial fields. Others, however, doubt its reliability and trustworthiness. This paper investigates the trustworthiness of ChatGPT and GPT-4 regarding logically consistent behaviour, focusing specifically on semantic consistency and the properties of negation, symmetric, and transitive consistency. Our findings suggest that while both models appear to show an enhanced language understanding and reasoning ability, they still frequently fall short of generating logically consistent predictions. We also ascertain via experiments that prompt designing, few-shot learning and employing larger large language models (LLMs) are unlikely to be the ultimate solution to resolve the inconsistency issue of LLMs.

Comments:	15 pages
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Journal reference:	The 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023)
Cite as:	arXiv:2303.06273 [cs.CL]
	(or arXiv:2303.06273v3 [cs.CL] for this version)

Submission history

From: Myeongjun Erik Jang [view email]
[v1] Sat, 11 Mar 2023 01:19:01 GMT (722kb,D)
[v2] Mon, 23 Oct 2023 20:40:39 GMT (888kb,D)
[v3] Tue, 14 Nov 2023 00:20:20 GMT (888kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2303.06273

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Consistency Analysis of ChatGPT

Submission history