Fusion of Domain Dependent Sensitive Semantics and Large Language Model for Research Text Security Classification Category
摘要
Research text security classification involves assigning appropriate security levels based on secrecy management regulations. Current approaches typically focus on text classification. However, due to specialized semantic dependencies within disciplinary fields, conventional methods often yield suboptimal results. Additionally, sensitive research texts, due to their highly confidential content, face challenges in acquiring annotated data and training models with limited samples. This paper proposes a fusion algorithm that integrates domain-specific sensitive semantics with large language model to address these challenges effectively. Experimental results demonstrate a 13% improvement in security classification accuracy compared to state-of-the-art techniques, highlighting its effectiveness in scenarios with limited samples. Further experiments on news text security classification tasks validate method’s applicability.