Opportunities to Use Arabic YouTube Comments in Text Mining and Masked Language Modeling
摘要
This study aims to extract meaningful information and explore some well-known Arabic-based transformer models in masked language modeling while investigating the interplay of the minimal length of the sentence and the percentage of masking. To this aim, we have used, for the first time, the opinions expressed in the comments of some trend YouTube channels in Morocco. By controlling the granularity of the input sequences during training, our findings indicate that employing 15 tokens as a minimal value can serve as an effective parameter during the fine-tuning of Arabic masked language models across the 15% and 40% masking rates.