A deep learning based method for identifying product manufacturing information in engineering drawings
摘要
Product Manufacturing Information (PMI) in engineering drawings (EDs) is critical for intelligent manufacturing, with its efficient and accurate extraction being pivotal to automating process design. This study proposes a hybrid framework integrating computer vision, deep learning, and multimodal large language model (MLLM) to tackle challenges in PMI extraction, such as scattered text distribution and complex specialized symbols. The framework implements a multi-stage processing workflow comprising object detection, rotation normalization, text recognition, and general information extraction. It innovatively employs oriented rotated box (OBB) detection and a four-class classification model to significantly improve PMI block identification accuracy. The research develops a hybrid text parsing architecture that seamlessly combines object detection and Optical Character Recognition (OCR), successfully extracting specialized content like tolerances and surface roughness. By leveraging structured prompts to drive the MLLM, the approach enables cross-format information extraction. Experimental results show that the proposed method outperforms existing technologies in PMI recognition accuracy and processing efficiency. Specifically, it achieves text recognition accuracies of 72% for local PMI, 87% for dimensional tolerance, 91% for geometric tolerance, and 88% for surface roughness, providing a comprehensive solution for automated PMI extraction in intelligent manufacturing.