A probabilistic similarity metric for Medline records: A model for author name disambiguation

A probabilistic similarity metric for Medline records: A model for author name disambiguation
复制标题

DOI:
10.1002/asi.20105
复制
发表时间:
2005-01-15
影响因子:
--
通讯作者:
Smalheiser, NR
Smalheiser, NR
中科院分区:
其他
文献类型:
--
作者:
Torvik, VI;Weeber, M;Smalheiser, NR

文献摘要

被引文献

相似文献

我们提出了一个模型来估计一对作者的名字(共享姓氏和首字母),出现在两个不同的Medline文章,指的是同一个人的概率。该模型使用一对文章之间简单而强大的相似性特征,基于标题,期刊名称,合著者姓名,医学主题词(MeSH),语言,从属关系和名称属性(文献中的流行程度,中间首字母和后缀)。相似性分布是从包含几乎完全作者匹配与不匹配的文章对的参考集计算的,以无偏的方式生成。虽然匹配集是自动生成的,并且可能包含一小部分不匹配,但该模型对不匹配的污染非常鲁棒。我们创建了一个免费的公共服务(“权威”:http://arrowsmith.psych.uic.edu),它将特定文章上的作者姓名作为输入,并将所有文章的列表作为输出,这些文章(姓氏,首字母)按相似性递减排列,并显示匹配概率。
We present a model for estimating the probability that a pair of author names (sharing last name and first initial), appearing on two different Medline articles, refer to the same individual. The model uses a simple yet powerful similarity profile between a pair of articles, based on title, journal name, coauthor names, medical subject headings (MeSH), language, affiliation, and name attributes (prevalence in the literature, middle initial, and suffix). The similarity profile distribution is computed from reference sets consisting of pairs of articles containing almost exclusively author matches versus nonmatches, generated in an unbiased manner. Although the match set is generated automatically and might contain a small proportion of nonmatches, the model is quite robust against contamination with nonmatches. We have created a free, public service ("Author-ity": http://arrowsmith.psych.uic.edu) that takes as input an author's name given on a specific article, and gives as output a list of all articles with that (last name, first initial) ranked by decreasing similarity, with match probability indicated.