To address the challenges posed by significant scale variations in scene text and the limitations of existing detection methods—particularly constrained receptive fields and inadequate multi-feature relationship modeling—this paper introduces a novel scene text detection framework integrating an improved feature pyramid with feature enhancement. The improved feature pyramid employs an adaptive feature fusion module to integrate shallow features, enhancing feature representation and enriching the model’s receptive field. The adaptive feature fusion module reduces information loss during feature fusion by optimizing relationships between multi-level features. The feature enhancement module extracts text contour features from multiple perspectives, improving adaptability to complex texts and compensating for the limitations of shallow networks in feature extraction. The experimental findings demonstrate the efficacy of the proposed approach, with F-measure reaching 83.7% on ICDAR 2015, 84.2% on CTW 1500, and 84.5% on MSRA-TD500.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Scene Text Detection Method with Improved Feature Pyramid and Feature Enhancement

  • Zhichao Xia,
  • Sheng Wang,
  • Tianhan Yang,
  • Xin Zhang

摘要

To address the challenges posed by significant scale variations in scene text and the limitations of existing detection methods—particularly constrained receptive fields and inadequate multi-feature relationship modeling—this paper introduces a novel scene text detection framework integrating an improved feature pyramid with feature enhancement. The improved feature pyramid employs an adaptive feature fusion module to integrate shallow features, enhancing feature representation and enriching the model’s receptive field. The adaptive feature fusion module reduces information loss during feature fusion by optimizing relationships between multi-level features. The feature enhancement module extracts text contour features from multiple perspectives, improving adaptability to complex texts and compensating for the limitations of shallow networks in feature extraction. The experimental findings demonstrate the efficacy of the proposed approach, with F-measure reaching 83.7% on ICDAR 2015, 84.2% on CTW 1500, and 84.5% on MSRA-TD500.