<p>Given the characteristics of aquaculture disease texts, which include the dense use of specialized terminology, diverse symptom descriptions, and complex disease progression stages, traditional classification methods often suffer from incomplete feature extraction and the inadequate utilization of domain knowledge. This study proposes a hybrid approach that combines deep learning with large model fine tuning, hereafter referred to as TABEM. The proposed approach utilizes an enhanced deep learning model for preliminary text classification and incorporates a dynamic adaptive multi-head attention (DAMHA) mechanism to effectively mitigate the loss of semantic information in long sequence texts. Additionally, a feature fusion module is integrated to further augment the model’s capacity to interpret complex texts. In cases of incorrect results, a large language model is employed for correction, and a fine-tuning dataset for this model is subsequently constructed. LoRA is employed to fine tune ChatGLM4-9B, thereby optimizing its performance within the aquaculture domain. Experimental results demonstrate that, through the collaborative interaction of the feature fusion module and the DAMHA mechanism, the model achieves an F1 score of 98.75%, surpassing other comparative models. The fine-tuned ChatGLM4-9B excels across BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L metrics, achieving scores of 99.52, 99.23, 99.24, and 99.46, respectively, thus demonstrating superior text classification accuracy. TABEM significantly enhances the fine-grained classification capability of aquaculture disease texts, thereby providing essential technological support for early disease warning and the development of precise diagnostic knowledge bases.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on fine-tuning strategies for text classification in the aquaculture domain by combining deep learning and large language models

  • Zhenglin Li,
  • Sijia Zhang,
  • Peirong Cao,
  • Jiaqi Zhang,
  • Zongshi An

摘要

Given the characteristics of aquaculture disease texts, which include the dense use of specialized terminology, diverse symptom descriptions, and complex disease progression stages, traditional classification methods often suffer from incomplete feature extraction and the inadequate utilization of domain knowledge. This study proposes a hybrid approach that combines deep learning with large model fine tuning, hereafter referred to as TABEM. The proposed approach utilizes an enhanced deep learning model for preliminary text classification and incorporates a dynamic adaptive multi-head attention (DAMHA) mechanism to effectively mitigate the loss of semantic information in long sequence texts. Additionally, a feature fusion module is integrated to further augment the model’s capacity to interpret complex texts. In cases of incorrect results, a large language model is employed for correction, and a fine-tuning dataset for this model is subsequently constructed. LoRA is employed to fine tune ChatGLM4-9B, thereby optimizing its performance within the aquaculture domain. Experimental results demonstrate that, through the collaborative interaction of the feature fusion module and the DAMHA mechanism, the model achieves an F1 score of 98.75%, surpassing other comparative models. The fine-tuned ChatGLM4-9B excels across BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L metrics, achieving scores of 99.52, 99.23, 99.24, and 99.46, respectively, thus demonstrating superior text classification accuracy. TABEM significantly enhances the fine-grained classification capability of aquaculture disease texts, thereby providing essential technological support for early disease warning and the development of precise diagnostic knowledge bases.