<p>With the rapid growth of mobile communication, spam SMS has become a major concern, particularly in multilingual contexts such as Indian regional languages, where detection remains challenging. While deep learning and large language models achieve promising performance, they demand large training datasets, contextual understanding of the messages, and computational resources. In contrast, machine learning algorithms are simpler and efficient but rely heavily on the quality of selected features, where traditional feature selection methods often struggle to identify the most significant subset of features, and it is crucial to utilize the strengths of these techniques. To address the challenges, this paper proposes a hand-crafted meta-data feature extraction approach for multilingual Dravidian spam SMS classification. The extracted features are evaluated using both a traditional feature selection method and a Graph-based Feature Evaluation Technique (GBFET) to validate their significance. Experiments on two Dravidian language datasets, namely RevisedIndianDataset and SpamSMSDataset, demonstrate that the proposed approach outperforms state-of-the-art methods with minimal features, offering an efficient solution for spam detection in low-resource languages.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient extraction and evaluation of hand-crafted meta-data features for Dravidian spam SMS classification

  • E. Ramanujam,
  • A. M. Abirami,
  • K. Sakthiprakash,
  • S. Sumitra

摘要

With the rapid growth of mobile communication, spam SMS has become a major concern, particularly in multilingual contexts such as Indian regional languages, where detection remains challenging. While deep learning and large language models achieve promising performance, they demand large training datasets, contextual understanding of the messages, and computational resources. In contrast, machine learning algorithms are simpler and efficient but rely heavily on the quality of selected features, where traditional feature selection methods often struggle to identify the most significant subset of features, and it is crucial to utilize the strengths of these techniques. To address the challenges, this paper proposes a hand-crafted meta-data feature extraction approach for multilingual Dravidian spam SMS classification. The extracted features are evaluated using both a traditional feature selection method and a Graph-based Feature Evaluation Technique (GBFET) to validate their significance. Experiments on two Dravidian language datasets, namely RevisedIndianDataset and SpamSMSDataset, demonstrate that the proposed approach outperforms state-of-the-art methods with minimal features, offering an efficient solution for spam detection in low-resource languages.