Minimization Techniques for Unstructured Data
摘要
Imagine having to find every instance of a person’s name across millions of emails, documents, images, and voice recordings—without missing a single occurrence. Now multiply that complexity with the over 50+ types of PII that might be associated with that individual. And now add the need to understand multilingual data. And don’t forget the disfluencies in writing, speech and transcription that you’re very likely going to have to deal with. The fundamental challenge of PII minimization for unstructured data is one of being able to identify the PII in the first place, a problem that has become exponentially more complex as organizations generate vast amounts of information in varied formats. While structured data’s predictable format makes it relatively straightforward to identify and protect personal information, as long as it has been stored in curated form and not haphazardly, unstructured data—which constitutes 80–90% of all enterprise data.