Dynamic taxonomy composition via keyqueries

Dynamic taxonomy composition via keyqueries
复制标题

DOI:
10.1109/jcdl.2014.6970148
复制
发表时间:
2014-09
期刊:
IEEE/ACM Joint Conference on Digital Libraries
影响因子:
--
通讯作者:
Tim Gollub;Michael Völske;Matthias Hagen;Benno Stein
Tim Gollub;Michael Völske;Matthias Hagen;Benno Stein
中科院分区:
其他
文献类型:
--
作者:
Tim Gollub;Michael Völske;Matthias Hagen;Benno Stein

文献摘要

相似文献

本文提出了一个无监督的框架,动态的,面向主题的分类组合在数字图书馆,它可以自然地集成现有的图书馆分类系统。在我们的方法中的分类类对应于所谓的关键字查询,是对数字图书馆的全文检索系统运行。给定一个文档,关键字查询是文档获得高相关性分数的几个关键字的集合。因此,关键字查询可以被看作是返回的检索结果的一般和简洁的描述。keyquery框架解决了静态分类系统的重要问题:过大的类和过于复杂的分类结构。例如,如果一个叶类增长到无法消化的大小,对所包含文档的键查询提供了一个合适的拆分机制。由于查询是众所周知的图书馆用户从他们的日常网络搜索经验,他们增加了结构的复杂性,在一个透明的方式。本文还提出了一种基于分类学的图书馆探索策略。鉴于用户的信息需要的形式库文档,我们合成一个层次结构的keyqueries,涵盖这个库的子集。我们设法解决这个困难的集覆盖问题上的飞行相结合的倒置和恢复索引沿着与启发式搜索空间修剪内的地图减少应用程序。ACM收集的科学论文的实证评估表明,我们的分类组成框架的效率和有效性。
This paper presents an unsupervised framework for dynamic, subject-oriented taxonomy composition in digital libraries, which can naturally integrate existing library classification systems. The taxonomy classes in our approach correspond to so-called keyqueries that are run against the digital library's full-text retrieval system. Given a document, a keyquery is a set of few keywords for which the document achieves a high relevance score. Keyqueries can hence be viewed as a general and concise description of the returned retrieval results. The keyquery framework addresses important problems of static classification systems: overlarge classes and overly complex taxonomy structures. If, for instance, a leaf class grows to an indigestible size, keyqueries for the contained documents provide a suitable split mechanism. Since queries are well-known to library users from their daily web search experience, they increase the structural complexity in a transparent way. The paper presents also a strategy for taxonomy-based library exploration. Given a user's information need in the form of library documents, we synthesize a hierarchy of keyqueries that covers this library subset. We manage to solve this difficult set covering problem on-the-fly by combining inverted and reverted indexes along with heuristic search space pruning within a map-reduce application. An empirical evaluation with an ACM collection of scientific papers demonstrates the efficiency and effectiveness of our taxonomy composition framework.