Learning to Create Customized Authority Lists

Learning to Create Customized Authority Lists
复制标题

DOI:
--
复制
发表时间:
2000-06
期刊:
--
影响因子:
--
通讯作者:
Huan Chang;David A. Cohn;A. McCallum
Huan Chang;David A. Cohn;A. McCallum
中科院分区:
其他
文献类型:
--
作者:
Huan Chang;David A. Cohn;A. McCallum

文献摘要

被引文献

相似文献

超文本的迅速发展和Kleinberg的HITS算法的流行引起了人们对链接分析的极大兴趣。HITS算法和文献计量学中较早的同类算法提供了一种确定某个主题的权威来源的方法,但它们不允许用户个人对什么来源是权威来源发表自己的看法。本文提出了一种学习用户内部权威模型的技术,目前的实验结果的基础上Cora在线索引数据库的约一百万在线计算机科学文献参考介绍文献计量学白色麦凯恩小涉及研究的结构,出现从一套链接的文件传统上,这些链接已采取的形式之间的引用期刊文章al虽然克莱因伯格和其他人,如布林佩奇已经发现,他们很好地适应了套超链接的文件文献计量学技术通过检查语料库中文档之间的引用或链接中的相关关系来识别主题领域、研究专业和对语料库的重要贡献。Kleinberg算法超文本诱导主题选择HITS计算文档链接矩阵的主特征向量,并根据文档在这些特征向量上的投影大小对文档的权威性进行排名。虽然最流行的、因此链接最多的文档在许多意义上可能是权威的,但它们可能不对应于特定用户的权威内部模型。用户可能有这样的看法,例如,特定文档X确实是权威的,或者作者Y所写的任何东西都是毫无价值的。在本文中,我们提出了一种技术,从少量的用户反馈中学习,以重新调整链接矩阵的特征向量,从而有效地重新校准权威度量,使其更接近用户的内部模型。我们的技术类似于相关反馈货车Rijsbergen,但不是操纵查询来学习检索到的文档的相关性,而是操纵链接矩阵的权重来学习我们首先演示如何使用该算法来提升某个特定子学科中研究论文的权威性,然后展示如何使用该算法自动创建用户指定的前十名列表。语料库遵循HITS算法的惯例然后我们描述了该方法在文档语料库中的应用,并讨论了我们在使用该方法时观察到的一些局限性。下面的部分描述了我们的提升算法,该算法解决了这些局限性中的一些。仅当存在从文档i到文档j的引用或链接时Mij通常被设置为文档所做的引用的总数或在某些变型中超过文档所做的引用的总数。对于集合中的给定文档i,设ai和hi分别为权威和中心分数。这些分数是大于或等于零的真实的数,并且具有以下解释:大的中心分数意味着文档指向许多好的权威,大的权威分数意味着它被许多好的枢纽所指向。这个递归定义导致了一组线性方程ai X
The proliferation of hypertext and the pop ularity of Kleinberg s HITS algorithm have brought about an increased interest in link analysis While HITS and its older relatives from the Bibliometrics provide a method for nding authoritative sources on a particular topic they do not allow individual users to inject their own opinions on what sources are authoritative This paper presents a tech nique for learning a user s internal model of authority We present experimental results based on Cora on line index a database of approximately one million on line computer science literature references Introduction Bibliometrics White McCain Small involves studying the structure that emerges from sets of linked documents Traditionally these links have taken the form of citations among journal articles al though Kleinberg and others e g Brin Page have found that they adapt well to sets of hyper linked documents Bibliometric techniques exist for identifying subject areas research specialties and in uential contributions to a corpus by examining cor relations in the references or links between documents in the corpus Kleinberg s algorithm Hypertext Induced Topic Se lection HITS calculates principal eigenvector of the document link matrix and ranks the authority of a document by the magnitude of its projection onto these eigenvectors A de ciency of this approach is that while the most popular and thus most heavily linked documents may be authoritative in many senses they may not correspond to a particular user s internal model of authority A user may have the opinion for example that a particular document X is truly an authority or that anything written by author Y is worthless What this user would like is a measure of document authority that coincides with their own preconceived notions of what is important In this paper we present a technique that learns from a small amount of user feedback to realign the eigen vectors of the link matrix e ectively re calibrating the measure of authority to correspond more closely to the user s internal model Our technique is similar to rel evance feedback van Rijsbergen but instead of manipulating a query to learn the relevance of the retrieved documents we manipulate the weighting of the link matrix to learn their authority We demonstrate our algorithm on the problem of iden tifying subjectively authoritative computer science re search papers We rst demonstrate how it can lift the authority of research papers in a particular sub discipline and then show how it can be used to auto matically create user speci c top ten lists The Authority of a Document In this section we rst describe how the authority of a document is computed with respect to a corpus following the conventions of the HITS algorithm We then describe the application of this method to a doc ument corpus and discuss some of the limitations we have observed with its use The following section de scribes our lifting algorithm which addresses some of these limitations Hypertext induced topic selection is performed as fol lows Given a set of documents a link matrixM spec i es their connectivity element Mij is non zero if and only if there is a reference or link from document i to document j TypicallyMij is set to or in some varia tions over the total number of references the document makes For a given document i in the set let ai and hi be the authority and hub scores respectively These scores are real numbers greater than or equal to zero and have the following interpretation a large hub score means the document points to many good authorities a large authority score means it is pointed to by many good hubs This recursive de nition leads to a set of linear equations ai X