Exploring advanced feature selection techniques: an application to dialectal Arabic data
摘要
In the field of automated language processing, distinguishing between Moroccan Arabic (Darija) in multilingual contexts is a major challenge. This study addresses this challenge by exploiting feature selection techniques to improve the accuracy of tongue detection. Using a comprehensive methodology integrating various feature selection methods, including TF-idf, CBOW, Word2Vec for feature extraction, LASSO regression, decision trees for machine learning techniques, and statistical feature selection such as ANOVA, Pearson correlation coefficient, and Mutual information, alongside semantic techniques by developing our semantic encoders adapted to Arabic data, we strive to uncover the most important, relevant language features while attenuating the noise. As a result, we have obtained robust results by applying the famous XGBOOST, with conventional extraction methods and SVD as a dimensional reduction method with a very reasonable execution time. Our experimental results highlight the effectiveness of feature selection techniques to reinforce supervised learning tasks. This research not only advances the field of natural language processing but also highlights the central role of feature selection in deciphering complex linguistic landscapes.