Papers
Topics
Authors
Recent
Search
2000 character limit reached

Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs

Published 22 Apr 2023 in cs.CL and cs.AI | (2304.11164v1)

Abstract: LLMs have become very popular recently and many claims have been made about their abilities, including for commonsense reasoning. Given the increasingly better results of current LLMs on previous static benchmarks for commonsense reasoning, we explore an alternative dialectical evaluation. The goal of this kind of evaluation is not to obtain an aggregate performance value but to find failures and map the boundaries of the system. Dialoguing with the system gives the opportunity to check for consistency and get more reassurance of these boundaries beyond anecdotal evidence. In this paper we conduct some qualitative investigations of this kind of evaluation for the particular case of spatial reasoning (which is a fundamental aspect of commonsense reasoning). We conclude with some suggestions for future work both to improve the capabilities of LLMs and to systematise this kind of dialectical evaluation.

Citations (20)

Summary

Paper to Video (Beta)

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.