BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing

Published 21 Jun 2022 in cs.CL | (2206.10668v2)

Abstract: Recent work has shown that generation from a prompted or fine-tuned LLM can perform well at semantic parsing when the output is constrained to be a valid semantic representation. We introduce BenchCLAMP, a Benchmark to evaluate Constrained LLM Parsing, that includes context-free grammars for seven semantic parsing datasets and two syntactic parsing datasets with varied output representations, as well as a constrained decoding interface to generate only valid outputs covered by these grammars. We provide low, medium, and high resource splits for each dataset, allowing accurate comparison of various LLMs under different data regimes. Our benchmark supports evaluation of LLMs using prompt-based learning as well as fine-tuning. We benchmark eight LLMs, including two GPT-3 variants available only through an API. Our experiments show that encoder-decoder pretrained LLMs can achieve similar performance or surpass state-of-the-art methods for syntactic and semantic parsing when the model output is constrained to be valid.