TH-DL Multilingual Text Recognition System Framework
摘要
By integrating a text detection model with text recognition models, the TH-DL text recognition system framework was designed and implemented for different tasks. Different text recognition models proposed in the book are compared. The experimental results show that PREN2D has achieved the highest recognition accuracy for word-level scene text images, whereas ARN-DTRN has achieved the best performance for sentence-level text images. A multilingual scene text recognition system based on the TH-DL framework was ranked first on the RRC-MLT-2019 leaderboard for the end-to-end text detection and recognition task. An Arabic video subtitle recognition system based on the TH-DL framework ranked first in the ICDAR 2017 and ICPR 2020 competitions on Arabic text detection and recognition in multiresolution video frames.