BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing

Roy, Subhro; Thomson, Sam; Chen, Tongfei; Shin, Richard; Pauls, Adam; Eisner, Jason; Van Durme, Benjamin

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2206

Change to browse by:

Computer Science > Computation and Language

Title: BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing

Authors: Subhro Roy, Sam Thomson, Tongfei Chen, Richard Shin, Adam Pauls, Jason Eisner, Benjamin Van Durme

(Submitted on 21 Jun 2022 (v1), last revised 10 Jan 2024 (this version, v2))

Abstract: Recent work has shown that generation from a prompted or fine-tuned language model can perform well at semantic parsing when the output is constrained to be a valid semantic representation. We introduce BenchCLAMP, a Benchmark to evaluate Constrained LAnguage Model Parsing, that includes context-free grammars for seven semantic parsing datasets and two syntactic parsing datasets with varied output representations, as well as a constrained decoding interface to generate only valid outputs covered by these grammars. We provide low, medium, and high resource splits for each dataset, allowing accurate comparison of various language models under different data regimes. Our benchmark supports evaluation of language models using prompt-based learning as well as fine-tuning. We benchmark eight language models, including two GPT-3 variants available only through an API. Our experiments show that encoder-decoder pretrained language models can achieve similar performance or surpass state-of-the-art methods for syntactic and semantic parsing when the model output is constrained to be valid.

Comments:	Neural Information Processing Systems (NeurIPS 2023) Track on Datasets and Benchmarks
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2206.10668 [cs.CL]
	(or arXiv:2206.10668v2 [cs.CL] for this version)

Submission history

From: Subhro Roy [view email]
[v1] Tue, 21 Jun 2022 18:34:11 GMT (62kb,D)
[v2] Wed, 10 Jan 2024 06:11:56 GMT (89kb)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2206.10668

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing

Submission history