GD-PTCF: Prompt-Tuning Based Classification Framework for Government Data
摘要
Government Data (GD), crucial for fostering social and economic growth, must adhere to specific classification standards and formats to ensure public accessibility and usability. Despite its potential, GD is currently hindered by a scarcity of high-quality, classified samples and the labor-intensive process of manual classification. To overcome these obstacles, our study introduces a Prompt-Tuning Classification Framework for Government Data (GD-PTCF), designed for automated classification. Initially, we employed web crawling techniques to amass an extensive dataset of Chinese government data. Subsequently, we unveiled a Classification Prompting Pattern (CPP) and utilized a BERT-based neural network, dubbed the Roberta Encoder (RE-coder), to facilitate few-shot prompt-tuning. This approach enables us to achieve remarkable classification accuracy with minimal training data. To further diminish the reliance on manual efforts, we developed a clustering mapping (CLM) strategy. This technique transforms encoded labeled embeddings into clustered embedding, which are then classified based on their proximity to predefined classification centers. Our experimental findings affirm that the GD-PTCF methodology significantly outperforms other pre-trained models in classification accuracy, even with a limited volume of training data.