Semantic Dependency Graph is a framework for representing deep semantic knowledge through flexible graph structures. While recent works indicate that large language models (LLMs) have impressive language and knowledge understanding abilities, it remains unclear whether they can understand this deep semantic knowledge. To explore this problem, we design four prompt-style probing tasks from aspects of semantic structure and semantic relations to adapt the inherent abilities of LLMs. To ensure thorough evaluation, we conduct extensive experiments in both in-context learning (ICL) and supervised fine-tuning (SFT) scenarios. Our findings indicate that the understanding of deep semantic knowledge requires larger parameter scale, especially the understanding of high-order semantic structure knowledge and semantic relation knowledge. Furthermore, our experiments reveal that while LLMs perform well on the in-domain (ID) test set via SFT, their generalization ability on out-of-domain (OOD) test set remains inadequate.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation and Analysis of the Chinese Semantic Dependency Understanding Ability of Large Language Models

  • Zizhuo Shen,
  • Wei Li,
  • Yanqiu Shao

摘要

Semantic Dependency Graph is a framework for representing deep semantic knowledge through flexible graph structures. While recent works indicate that large language models (LLMs) have impressive language and knowledge understanding abilities, it remains unclear whether they can understand this deep semantic knowledge. To explore this problem, we design four prompt-style probing tasks from aspects of semantic structure and semantic relations to adapt the inherent abilities of LLMs. To ensure thorough evaluation, we conduct extensive experiments in both in-context learning (ICL) and supervised fine-tuning (SFT) scenarios. Our findings indicate that the understanding of deep semantic knowledge requires larger parameter scale, especially the understanding of high-order semantic structure knowledge and semantic relation knowledge. Furthermore, our experiments reveal that while LLMs perform well on the in-domain (ID) test set via SFT, their generalization ability on out-of-domain (OOD) test set remains inadequate.