A bag of features (BOF) may be made using either map reduction techniques or a combination of a thesaurus and domain knowledge. This research presents the BOFMR (Bag of Features using MapReduce) and BOFWT (Bag of Features with Weighted Terms) algorithms, a scalable and efficient technique for processing large email datasets and generating feature vectors based on pre-defined characteristics. The outcomes from using both BOFs on identical datasets are compared. The algorithm leverages the parallel processing capabilities of the MapReduce framework to handle the extensive data, ensuring performance and scalability. When creating a bag of words from a training dataset, the BOFMR technique is useful. The map-reduce technique will help to create a bag of features faster even in case of a larger chunk of data. In this experiment, as data size was limited, the performance of map reduce was not measured. In another BOFWT approach, the building of BOF with domain knowledge by using the word thesaurus was a challenge. The experimental result shows that the results of BOFWT are nearer to the output of BOFMR, and both algorithms show the highest accuracy among other methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Phishing Email Classification Through Scalable Feature Extraction Using MapReduce

  • Syed Hameed Uddin,
  • Mugaerah Ahmed Shareef Maaz,
  • Himanshu Gupta,
  • Arya Jinan Panicker,
  • Darshan Upadhyay,
  • Shravan Khunti,
  • Kamal Upreti

摘要

A bag of features (BOF) may be made using either map reduction techniques or a combination of a thesaurus and domain knowledge. This research presents the BOFMR (Bag of Features using MapReduce) and BOFWT (Bag of Features with Weighted Terms) algorithms, a scalable and efficient technique for processing large email datasets and generating feature vectors based on pre-defined characteristics. The outcomes from using both BOFs on identical datasets are compared. The algorithm leverages the parallel processing capabilities of the MapReduce framework to handle the extensive data, ensuring performance and scalability. When creating a bag of words from a training dataset, the BOFMR technique is useful. The map-reduce technique will help to create a bag of features faster even in case of a larger chunk of data. In this experiment, as data size was limited, the performance of map reduce was not measured. In another BOFWT approach, the building of BOF with domain knowledge by using the word thesaurus was a challenge. The experimental result shows that the results of BOFWT are nearer to the output of BOFMR, and both algorithms show the highest accuracy among other methods.