错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Can Large Language Models Put 2 and 2 Together? Probing for Entailed Arithmetical Relationships

  • Dagmara Panas,
  • Sohan Seth,
  • Vaishak Belle

摘要

Two major areas of interest in the era of Large Language Models regard questions of what do LLMs know, and if and how they may be able to reason, or rather, approximately reason. Since to date these lines of work progressed largely in parallel, we are interested in investigating the intersection: probing for reasoning about the implicitly-held knowledge. Suspecting the performance to be lacking, we use a very simple set-up of comparisons between cardinalities associated with elements of various subjects (e.g. the number of legs of a bird vs. the number of wheels on a tricycle). We empirically demonstrate that although LLMs make steady progress in knowledge acquisition and (pseudo)reasoning with each new GPT release, they remain limited in their capabilities to performing probabilistic retrieval. We argue that pure statistical learning can not cope with the combinatorial explosion inherent in many commonsense reasoning tasks. Further, we emphasise that bigger is not always better and chasing purely statistical improvements is flawed at the core, since it only exacerbates the dangerous conflation of the production of correct answers with genuine reasoning ability.