Transformers and Spark for Automated CV Classification in Arabophone Regions
摘要
The growing demand for skilled labor in Arabophone regions has led to increased interest in automating the Curriculum Vitae (CV) classifying process. In this study, we present an innovative methodology that exploits the power of Transformers models and the flexibility of the Spark platform for the automatic classification of CVs written in Arabic. We start by preprocessing CVs using Arabic-specific text processing techniques, including tokenization and normalization. Then, we use a pre-trained Transformer model, adapted to the Arabic linguistic context, to extract relevant features from the CVs. These features are then fed into a Spark pipeline for classification. Thanks to the scaling of Spark, we were able to process large quantities of CVs in record time, making it a practical solution for recruitment companies. This research paves the way for more efficient and accurate automation of Arabic CV classifying, helping to facilitate the recruitment process in Arabic-speaking regions.