<p>This paper presents a new feature extraction technique for Optical Character Recognition (OCR) that achieves state-of-the-art results in the recognition of multilingual characters, overcoming scale, rotation, and distortion challenges. Our method leverages polar coordinates to incorporate two innovative features: Extended Ellipse-Based Features (EEB) and Crossing Count Measure (CCM), which provide inherent scale and rotational invariance. From the benchmark datasets, the proposed technique was tested on character datasets such as ISI Bengali and Chars74K, yielding accuracy rates of 98.82% and 98.69% respectively. These statistics also depict high precision and recall values coupled with a high F1-score, indicating that the method has been sound and robust. Interestingly, our approach exhibits far less computational overhead compared to traditional CNN-based methods and is, therefore, a good candidate to be deployed on resource-constrained edge devices. This work fills the gap between high-performance OCR systems and practical deployment needs by providing a scalable and efficient solution for multilingual character recognition in diverse and challenging contexts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An effective geometrical feature extraction method for scale and rotational invariant multi-lingual character recognition

  • Sharfuddin Waseem Mohammed,
  • Brindha Murugan

摘要

This paper presents a new feature extraction technique for Optical Character Recognition (OCR) that achieves state-of-the-art results in the recognition of multilingual characters, overcoming scale, rotation, and distortion challenges. Our method leverages polar coordinates to incorporate two innovative features: Extended Ellipse-Based Features (EEB) and Crossing Count Measure (CCM), which provide inherent scale and rotational invariance. From the benchmark datasets, the proposed technique was tested on character datasets such as ISI Bengali and Chars74K, yielding accuracy rates of 98.82% and 98.69% respectively. These statistics also depict high precision and recall values coupled with a high F1-score, indicating that the method has been sound and robust. Interestingly, our approach exhibits far less computational overhead compared to traditional CNN-based methods and is, therefore, a good candidate to be deployed on resource-constrained edge devices. This work fills the gap between high-performance OCR systems and practical deployment needs by providing a scalable and efficient solution for multilingual character recognition in diverse and challenging contexts.