利用姓名关联个人记录

The Use of Names for Linking Personal Records

Journal of the American Statistical Association · 1992
被引 8
ABS 4

中文导读

研究了计算机如何像人类一样利用姓名变体(如ANTHONY与TONY)的记忆来搜索和关联个人记录,并量化了六种姓名匹配规则的改进效果,其中计算同义词姓名比较的ODDS值最为关键。

Abstract

Abstract The skill of a human who searches large files of personal records depends much on prior knowledge of how the names vary in successive documents pertaining to the same individuals (e.g., as with ANTHONY–TONY, JOSEPH–JOE, WILLIAM–BILL). Now, an essentially exact procedure enables computers to make similar use of an accumulated memory of their own past experiences when searching for, and linking, records that relate to particular persons. This knowledge is further applied to quantify the benefits from various refinements of the rules by which the discriminating powers of names are calculated when they do not precisely agree or are substantially dissimilar. Of the six refinements tested, by far the most important is the recently developed exact approach for calculating the ODDS associated with comparisons of names that are possible synonyms.

信息检索机器学习逻辑回归计算机科学