Name Pattern Recognition: A Model Proposal Applied to the Anonymization of Unstructured Data
摘要
The increasing digitalization of information has heightened concerns about the privacy of personal data, especially in unstructured textual documents. In this context, anonymization emerges as a crucial tool to ensure compliance with data protection regulations, such as the General Data Protection Law (LGPD) in Brazil. This paper presents a solution for the identification and anonymization of personal names in texts in Brazilian Portuguese, focusing on police reports (BOs), using natural language processing and machine learning techniques. Our approach employs a computational model that combines different machine learning techniques, integrating grammatical classification and context analysis to refine the classification of personal names and minimize the occurrence of false negatives. The results obtained were compared with BERTimbau and spaCy, reference models in the field. The methodology was evaluated using BOs provided by the Civil Police of Pará, achieving a recall of 99.62% in identifying personal names. While the increase in recall contributed to a rise in the number of false positives, this trade-off was necessary to maximize the coverage of personal names in the anonymization process. Furthermore, entropy analysis demonstrated that anonymization preserved textual variability, supporting the robustness of the proposal. The research results reinforce the methodology’s potential to enhance privacy protection in unstructured texts, ensuring compliance with regulations and strengthening data security in sensitive contexts.