Recently launched large language models (LLMs), such as ChatGPT, Copilot, and Gemini, have demonstrated impressive capabilities in understanding and responding to human queries across various domains. However, thorough evaluation is needed, as human language involves ambiguities at multiple levels-phonetic, lexical, syntactic, pragmatic, and discourse. These ambiguities are not merely inherent but contribute to communicative efficiency, posing significant challenges for LLMs. To address this, 24 pragmatic and discourse ambiguous sentences were curated and tested across ChatGPT 4.0, Copilot, and Gemini-an area often overlooked in studies on Generative AI tools due to its contextual complexity. The results indicate that ambiguity is not just an inherent feature but a key element of human language, enhancing communication efficiency through nuanced contextual understanding-a phenomenon still difficult for AI tools to fully comprehend. Notably, ChatGPT showed better results than other tools, achieving 75% accuracy. These findings highlight the need to integrate more nuanced linguistic rules and frameworks into LLMs to improve contextual interpretation and ambiguity handling.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Language of Nuance: Exploring the Limits of Large Language Models in Handling Ambiguity

  • Md. Tauseef Qamar,
  • Shahab Saquib Sohail,
  • Gunjan Ansari,
  • Chandni Saxena

摘要

Recently launched large language models (LLMs), such as ChatGPT, Copilot, and Gemini, have demonstrated impressive capabilities in understanding and responding to human queries across various domains. However, thorough evaluation is needed, as human language involves ambiguities at multiple levels-phonetic, lexical, syntactic, pragmatic, and discourse. These ambiguities are not merely inherent but contribute to communicative efficiency, posing significant challenges for LLMs. To address this, 24 pragmatic and discourse ambiguous sentences were curated and tested across ChatGPT 4.0, Copilot, and Gemini-an area often overlooked in studies on Generative AI tools due to its contextual complexity. The results indicate that ambiguity is not just an inherent feature but a key element of human language, enhancing communication efficiency through nuanced contextual understanding-a phenomenon still difficult for AI tools to fully comprehend. Notably, ChatGPT showed better results than other tools, achieving 75% accuracy. These findings highlight the need to integrate more nuanced linguistic rules and frameworks into LLMs to improve contextual interpretation and ambiguity handling.