Using Linkage Context for Automated Correction in Unsupervised Entity Resolution
摘要
Many modern Entity Resolution (ER) systems leverage metadata about the reference data to facilitate processing and making equivalence decisions. This historically has required that each source of input data be pre-processed individually to conform to a common metadata alignment, have data cleansing applied, and the data condensed into a singular dataset to be submitted to an ER process. These are costly processes and require additional passes of the input data prior to ER decisioning. This paper expands on the concept of context within ER to replace the need for metadata alignment and preprocessing that was introduced in previous literature. Leveraging the four (4) points within ER processes for which context can be extracted, and introducing a new n-gram similarity method allows for better more intelligent corrections to be performed to the data during processing without the need for preprocessing or metadata. This paper defines tested methods to enhance the initial context extraction to reduce erroneous automated corrections that thusly impact the final clustering results of the ER system and provides empirical results to support the effectiveness of these methods.