Scene text detection has been a vital region of inquire about in computer vision for an expanded period. Recognizing text in natural scenes is essential for numerous applications. However, this task is not straightforward. There are complex backgrounds and substantial variations in fonts, sizes, colors, and orientations. This study introduces a novel approach to recognizing text in scenes, employing advanced machine learning techniques. By combining deep learning with the latest image processing techniques, this system enhances both text detection and recognition capabilities. It performs well across diverse and dynamic scenes. Traditional methods struggle, particularly with Devanagari script feature extraction and handling irregular backgrounds. This research presents three unique neural network models: Hybrid Convolutional Neural Network (CNN), Fusion Neural Network (FNN), and Ensemble Deep Network (EDN). Initially, the study develops a robust method for detecting text in images from natural scenes. This includes constructing a “CRF model” that connects “CNN scores of Maximally Stable External Regions (MSERs)” with various neighborhood information. The research also utilizes YOLO-based object detection alongside CNN-based classification. In the second part, it focuses on the FNN model, which combines Convolutional and Recurrent Neural Networks. Here, convolutional layers extract features, while recurrent layers improve the prediction of sequences of these features, leading to enhanced accuracy in classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Study on Framework for Integrated Text Detection and Recognition in Natural Image’s Through Unified Deep Learning Approaches

  • Pallavi Krishna Purohit,
  • Vikas Somani,
  • Sandeep Saxena

摘要

Scene text detection has been a vital region of inquire about in computer vision for an expanded period. Recognizing text in natural scenes is essential for numerous applications. However, this task is not straightforward. There are complex backgrounds and substantial variations in fonts, sizes, colors, and orientations. This study introduces a novel approach to recognizing text in scenes, employing advanced machine learning techniques. By combining deep learning with the latest image processing techniques, this system enhances both text detection and recognition capabilities. It performs well across diverse and dynamic scenes. Traditional methods struggle, particularly with Devanagari script feature extraction and handling irregular backgrounds. This research presents three unique neural network models: Hybrid Convolutional Neural Network (CNN), Fusion Neural Network (FNN), and Ensemble Deep Network (EDN). Initially, the study develops a robust method for detecting text in images from natural scenes. This includes constructing a “CRF model” that connects “CNN scores of Maximally Stable External Regions (MSERs)” with various neighborhood information. The research also utilizes YOLO-based object detection alongside CNN-based classification. In the second part, it focuses on the FNN model, which combines Convolutional and Recurrent Neural Networks. Here, convolutional layers extract features, while recurrent layers improve the prediction of sequences of these features, leading to enhanced accuracy in classification.