Leveraging Large Language Models for QA Dialogue Dataset Construction and Analysis in Public Services
摘要
This paper identifies the limitations of current AI datasets within the public service sector, specifically concerning the human-robot interaction (HRI) context. Existing datasets often lack the necessary interactive features for effective and efficient interactions, hindering the development of customized and emotionally responsive systems. As public service demands become more diverse and complex in HRI, traditional datasets fail to support high-quality interactions, necessitating significant improvements. To address this issue, we introduce a QA dialogue dataset specifically tailored for public service applications, comprising 1208 pairs generated by large language model. This dataset integrates textual and emotional data, providing detailed annotations for interaction quality and emotional accuracy. Our method includes four stages: data generation, annotation, emotion analysis, and performance evaluation. During the data generation stage, GPT-4 is employed to create a diverse set of dialogues. In the annotation stage, these dialogues are meticulously labeled for quality and emotional content. The emotion analysis stage utilizes various recognition algorithms to process the data. Finally, the performance evaluation stage involves experiments to validate the dataset’s effectiveness. Comparative experiments demonstrate the dataset’s efficacy in enhancing the adaptability and performance of public service robots, underscoring its potential for training AI models to effectively handle real-world dialogues.