Integrating Cross-Language Hierarchies and Its Application to Retrieving Relevant Documents

Integrating Cross-Language Hierarchies and Its Application to Retrieving Relevant Documents
复制标题

DOI:
10.1145/1386869.1386870
复制
发表时间:
2008-06
期刊:
ACM Trans. Asian Lang. Inf. Process.
影响因子:
--
通讯作者:
Fumiyo Fukumoto;Yoshimi Suzuki
Fumiyo Fukumoto;Yoshimi Suzuki
中科院分区:
其他
文献类型:
--
作者:
Fumiyo Fukumoto;Yoshimi Suzuki

文献摘要

相似文献

互联网目录,如Yahoo!是一种提高Web上信息检索(IR)的有效性和效率的方法,因为页面(文档)被组织成分层的类别,并且相似的页面被分组在一起。Web服务上的大多数搜索引擎查找被分配到单一分类层次结构的文档。层次结构中的类别由人类专家仔细定义,文件组织良好。然而,一种语言的单一层次往往不足以找到所有相关材料,因为每个层次在定义层次结构和对文档进行分类方面往往都有一些偏见。此外,以用户母语以外的语言编写的文档通常包括与用户请求相关的大量信息。在本文中,我们提出了一种通过估计类别相似度来整合日语的跨语言(CL)类别层次,即路透社96层次和UDC代码层次的方法。该方法不是简单地将两个不同的层次合并成一个大的层次,而是提取相似类别的集合,其中集合的每个元素都相互关联。它由三个步骤组成。首先,我们使用跨语言文本分类(CLTC)技术将文档从一个层次分类到另一个层次的类别,并提取两个层次的类别对。然后,对这些类别对进行统计,得到相似的类别对;最后,对这些类别对,应用关联规则算法的生成函数,找出相似的类别集。此外,我们还研究了整合层次结构是否有助于支持检索内容相似的文档。检索结果表明,与基本无层次模型相比,改进了42.7%,比单一层次模型提高了21.6%。
Internet directories such as Yahoo! are an approach to improvethe efficacy and efficiency of Information Retrieval (IR) on theWeb, as pages (documents) are organized into hierarchicalcategories, and similar pages are grouped together. Most of thesearch engines on the Web service find documents that are assignedto a single classification hierarchy. Categories in the hierarchyare carefully defined by human experts and documents are wellorganized. However, a single hierarchy in one language is ofteninsufficient to find all relevant material, as each hierarchy tendsto have some bias in both defining hierarchical structure andclassifying documents. Moreover, documents written in a languageother than the users native language often include large amounts ofinformation related to the users request. In this article, wepropose a method of integrating cross-language (CL) categoryhierarchies, that is, Reuters 96 hierarchy and UDC code hierarchyof Japanese by estimating category similarities. The method doesnot simply merge two different hierarchies into one large hierarchybut instead extracts sets of similar categories, where each elementof the sets is relevant with each other. It consists of threesteps. First, we classify documents from one hierarchy intocategories with another hierarchy using a cross-language textclassification (CLTC) technique, and extract category pairs of twohierarchies. Next, we apply Ç2 statisticsto these pairs to obtain similar category pairs, and finally weapply the generating function of the Apriori algorithm(Apriori-Gen) to the category pairs, and find sets of similarcategories. Moreover, we examined whether integrating hierarchieshelps to support retrieval of documents with similar contents. Theretrieval results showed a 42.7% improvement over the baselinenonhierarchy model, and a 21.6% improvement over a singlehierarchy.