<p>Probabilistic record linkage is often used to match records from two files, in particular when the variables common to both files comprise identifiers measured with occasional errors like names and demographic variables. We consider bipartite record linkage settings in which there are no duplicates within the files, but some entities appear in both files. In this setting, the analyst desires a point estimate of the linkage structure that matches each record to at most one record from the other file. We introduce an estimator that maximizes the expected <i>F</i>-score for the linkage structure. We target the approach for record linkage methods that produce either (an approximate) posterior distribution of the unknown linkage structure or probabilities of matches for record pairs. Using simulations and applications with genuine data, we illustrate that the <i>F</i>-score estimators have desirable properties.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimal F-score Matching for Bipartite Record Linkage

  • Eric A. Bai,
  • Olivier Binette,
  • Jerome P. Reiter

摘要

Probabilistic record linkage is often used to match records from two files, in particular when the variables common to both files comprise identifiers measured with occasional errors like names and demographic variables. We consider bipartite record linkage settings in which there are no duplicates within the files, but some entities appear in both files. In this setting, the analyst desires a point estimate of the linkage structure that matches each record to at most one record from the other file. We introduce an estimator that maximizes the expected F-score for the linkage structure. We target the approach for record linkage methods that produce either (an approximate) posterior distribution of the unknown linkage structure or probabilities of matches for record pairs. Using simulations and applications with genuine data, we illustrate that the F-score estimators have desirable properties.