<p>Disease prevention and control is crucial for the healthy development of the aquaculture industry, and high-quality data are the foundation for intelligent disease management. However, textual data in this field are scarce and variable in quality. The direct use of large language models (LLMs) for data augmentation often results in poor-quality outputs. In this paper, we propose a small-sample data augmentation framework called TDA-GLM. This framework employs a strategy that combines a high-quality corpus, pretrained large model, and efficient prompts, which improves data augmentation by integrating high-quality corpora and refining prompt techniques. To address the poor adaptability of LLMs in aquaculture disease prevention and control, which leads to low-quality data augmentation, we introduce a supervised learning model. This model guides the understanding of domain knowledge and optimizes data augmentation tasks within the LLM, enabling the generation of data that are both highly similar to the original professional material and rich in content. Additionally, we design a noise removal module that filters out noise by analyzing the consistency between the augmented data and the data augmentation target, thereby increasing the overall quality of the augmented data. To verify the domain reliability of our data augmentation approach, we conducted a few-shot learning classification task experiment. The results demonstrated significant improvements by our model over existing advanced text data augmentation techniques, with the key performance indicators Acc, P, R, and F1 reaching 94.86%, 95.55%, 94.86%, and 94.62%, respectively. Furthermore, in the experiments assessing the quality of the augmented samples, the similarity and enrichment degrees reached 94.95% and 64.41%, respectively. These results indicate that our augmented samples are of very high quality, ensuring that the core semantics of the original data are preserved while increasing data diversity through appropriate variations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TDA-GLM: text data augmentation for aquaculture disease prevention and control via a small model-guided ChatGLM

  • Shihan Qiao,
  • Hong Yu,
  • Jing Song,
  • Lixin Zhang,
  • Huiyuan Zhao,
  • Wei Huang

摘要

Disease prevention and control is crucial for the healthy development of the aquaculture industry, and high-quality data are the foundation for intelligent disease management. However, textual data in this field are scarce and variable in quality. The direct use of large language models (LLMs) for data augmentation often results in poor-quality outputs. In this paper, we propose a small-sample data augmentation framework called TDA-GLM. This framework employs a strategy that combines a high-quality corpus, pretrained large model, and efficient prompts, which improves data augmentation by integrating high-quality corpora and refining prompt techniques. To address the poor adaptability of LLMs in aquaculture disease prevention and control, which leads to low-quality data augmentation, we introduce a supervised learning model. This model guides the understanding of domain knowledge and optimizes data augmentation tasks within the LLM, enabling the generation of data that are both highly similar to the original professional material and rich in content. Additionally, we design a noise removal module that filters out noise by analyzing the consistency between the augmented data and the data augmentation target, thereby increasing the overall quality of the augmented data. To verify the domain reliability of our data augmentation approach, we conducted a few-shot learning classification task experiment. The results demonstrated significant improvements by our model over existing advanced text data augmentation techniques, with the key performance indicators Acc, P, R, and F1 reaching 94.86%, 95.55%, 94.86%, and 94.62%, respectively. Furthermore, in the experiments assessing the quality of the augmented samples, the similarity and enrichment degrees reached 94.95% and 64.41%, respectively. These results indicate that our augmented samples are of very high quality, ensuring that the core semantics of the original data are preserved while increasing data diversity through appropriate variations.