Phishing Text Email Detection by Natural Language Processing with Machine Learning
摘要
Bag of features may be made using either map reduction techniques or a combination of a thesaurus and domain knowledge. The outcomes from using both BOF on the identical datasets are compared. When creating a bag of words from a training dataset, the BOFMR technique is useful. The map reduce technique will help to create bag of features faster even in case of larger chunk of data. In this experiment as data size was limited, we have not measured the performance of map reduce. In another BOFWT approach, the building of BOF with domain knowledge by using word thesaurus was challenge. The experimental result shows that the results of BOFWT are nearer to the output of BOFMR.