BRIDP: Dataset and Validation Method for BRazilian Identity Document Parsing
摘要
The manual processing of identity documents (ID) introduces various challenges, including the potential for human errors, low efficiency, and a lack of standardization. Automation emerges as the solution to overcome these obstacles, offering enhanced precision, speed, and consistency in the document validation process. In this context, we introduce a method for validating three types of Brazilian identity documents: the National Driver’s License (CNH), the Individual Taxpayer Registry (CPF), and the General Registry (RG). Our approach involves the extraction of pertinent information from each document type using cutting-edge convolutional neural networks (CNN) for object detection, followed by optical character recognition (OCR). To facilitate our research and contribute to the identity document parsing field of study, we created a set of synthetic documents for model training and validation. The offered dataset is freely available to researchers, and the results obtained, varying from 65% to 91% accuracy, affirm that the proposed method is promising for the validation of identity documents in the digital age.