A comprehensive Data Resource System could provide semantic standardization across data from various sources, facilitating cross-business recognition and serving as the cornerstone in addressing the challenge of data silos. However, existing systems, e.g., the Urban Knowledge System (UKS), primarily focus on offering data abstractions, still challenging in refining the semantics of data. These difficulties arise from the openness and the redundancy of textual semantics, which can result in differences among a nuanced distinction of textual terms. To address these challenges, we propose the Term Retrieval Augmented Generation (TermRAG) to precisely refine the semantics of data and expand the terms. Specifically, TermRAG extracts the correlated expressions from historical data, and proactively expands them with foresight by the Large Language Models (LLMs). With the alignment of the human understanding and the generating processes, it greatly improves the generated quality and provides crucial insights to experts. We benchmark our approach against other baselines under three real-world datasets, with results highlighting our method’s superiority. The refined terms have been successfully integrated into Beijing’s municipal government applications, significantly boosting data interoperability efficiency.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TermRAG: Data Resource Summarization via Term Retrieval Augmented Generation

  • Chunyang Li,
  • Yanping Sun,
  • Zhichao Huang,
  • Jinjin Guo

摘要

A comprehensive Data Resource System could provide semantic standardization across data from various sources, facilitating cross-business recognition and serving as the cornerstone in addressing the challenge of data silos. However, existing systems, e.g., the Urban Knowledge System (UKS), primarily focus on offering data abstractions, still challenging in refining the semantics of data. These difficulties arise from the openness and the redundancy of textual semantics, which can result in differences among a nuanced distinction of textual terms. To address these challenges, we propose the Term Retrieval Augmented Generation (TermRAG) to precisely refine the semantics of data and expand the terms. Specifically, TermRAG extracts the correlated expressions from historical data, and proactively expands them with foresight by the Large Language Models (LLMs). With the alignment of the human understanding and the generating processes, it greatly improves the generated quality and provides crucial insights to experts. We benchmark our approach against other baselines under three real-world datasets, with results highlighting our method’s superiority. The refined terms have been successfully integrated into Beijing’s municipal government applications, significantly boosting data interoperability efficiency.