Using an Optimal then Enhanced YOLO Model for Multi-Lingual Scene Text Detection Containing the Arabic Scripts
摘要
In recent years, significant advancements have been made in deep learning and the recognition of text in images of natural scenes, thanks to the advancements in machine learning and artificial intelligence. The limited availability of diverse datasets containing multiple languages and scripts often restricts the effectiveness of deep learning and text detection in the wild, particularly when it comes to Arabic language as an additional challenge. Despite notable progress, this scarcity remains a constraint. The deep learning neural network known as YOLO (You Only Look Once) has become widely popular due to its versatility in addressing a wide range of machine learning tasks, particularly in the domain of computer vision. The YOLO algorithm has gained increasing acknowledgment for its outstanding ability to tackle complex problems in conjunction with complex backgrounds of an image captured from nature, handle noisy data, and overcome various challenges encountered in real-world situations. Our experiments offer a succinct analysis of text detection algorithms that rely on convolutional neural networks (CNNs); In particular, we focus on various iterations of the YOLO models, employing same specific data augmentation techniques on both SYPHAX dataset and ICDAR MLT-2019 dataset, which comprise Arabic scripts in real natural scene images. The aim of this article is to identify the most effective YOLO algorithm for detecting text containing the Arabic scripts in the wild then to enhance this optimal model obtained in addition to explore potential research avenues that can enhance the capabilities of the most robust architecture in this field.