Disambiguating Web appearances of people in a social network

Disambiguating Web appearances of people in a social network
复制标题

DOI:
10.1145/1060745.1060813
复制
发表时间:
2005-05
影响因子:
5.4
通讯作者:
Ron Bekkerman;A. McCallum
Ron Bekkerman;A. McCallum
中科院分区:
医学2区
文献类型:
--
作者:
Ron Bekkerman;A. McCallum

文献摘要

被引文献

相似文献

假设你正在寻找关于某个特定人的信息。搜索引擎会为这个人的名字返回很多页面,但哪些页面是关于你关心的人的,哪些是关于碰巧同名的其他人的?此外,如果我们正在寻找在某种程度上有关系的多个人,我们如何才能最好地利用这个社交网络?本文提出了两个无监督框架来解决这个问题:一个基于Web页面的链接结构,另一个使用凝聚/凝聚双重聚类(A/CDC)--这是最近引入的一种多向分布式聚类方法的应用。为了评估我们的方法,我们收集并手工标记了1000多个网页的数据集,这些网页是从谷歌查询中检索到的,这些网页上有12个人的名字一起出现在电子邮件文件夹中的某个人身上。在此数据集上,我们的方法比传统的凝聚聚类性能高出20%以上,达到了80%以上的F-MEASURE。
Say you are looking for information about a particular person. A search engine returns many pages for that person's name but which pages are about the person you care about, and which are about other people who happen to have the same name? Furthermore, if we are looking for multiple people who are related in some way, how can we best leverage this social network? This paper presents two unsupervised frameworks for solving this problem: one based on link structure of the Web pages, another using Agglomerative/Conglomerative Double Clustering (A/CDC)---an application of a recently introduced multi-way distributional clustering method. To evaluate our methods, we collected and hand-labeled a dataset of over 1000 Web pages retrieved from Google queries on 12 personal names appearing together in someones in an email folder. On this dataset our methods outperform traditional agglomerative clustering by more than 20%, achieving over 80% F-measure.