Descriptor-Invariant Fusion Architectures for Automatic Subject Indexing

Descriptor-Invariant Fusion Architectures for Automatic Subject Indexing
复制标题

用于自动主题索引的描述符不变融合架构

DOI:
--
复制
发表时间:
2017
期刊:
ACM/IEEE Joint Conference on Digital Libraries
影响因子:
--
通讯作者:
C. Seifert
C. Seifert
中科院分区:
--
文献类型:
--
作者:
Martin Toepfer;C. Seifert

文献摘要

被引文献

相似文献

使用受控词汇索引的文档使图书馆用户能够发现相关文档,甚至跨越语言障碍。由于科学出版物的快速增长,数字图书馆需要自动的方法,准确地索引文件,特别是关于显式或隐式的概念漂移,即,关于新的描述符术语和新类型的文件,分别。本文首先分析了自动标引相关方法的体系结构。我们表明,他们的设计决定了个人的优势和劣势,并证明他们的融合研究。特别是,系统受益于统计关联组件以及应用字典匹配,排名和二进制分类的词汇组件。分析强调了特征不变量学习的重要性,即基于特征的学习可以在不同的描述符之间转移。经济标题和作者关键词的理论和实验结果强调的相关性的融合方法的整体准确性和适应性的动态域。实验表明,融合策略相结合的二进制相关性的方法和基于词库的系统优于所有其他策略的测试数据集。本文的研究结果可以帮助数字图书馆的研究人员和实践者选择合适的自动标引方法。
Documents indexed with controlled vocabularies enable users of libraries to discover relevant documents, even across language barriers. Due to the rapid growth of scientific publications, digital libraries require automatic methods that index documents accurately, especially with regard to explicit or implicit concept drift, that is, with respect to new descriptor terms and new types of documents, respectively. This paper first analyzes architectures of related approaches on automatic indexing. We show that their design determines individual strengths and weaknesses and justify research on their fusion. In particular, systems benefit from statistical associative components as well as from lexical components applying dictionary matching, ranking, and binary classification. The analysis emphasizes the importance of descriptor-invariant learning, that is, learning based on features which can be transferred between different descriptors. Theoretic and experimental results on economic titles and author keywords underline the relevance of the fusion methodology in terms of overall accuracy and adaptability to dynamic domains. Experiments show that fusion strategies combining a binary relevance approach and a thesaurus-based system outperform all other strategies on the tested data set. Our findings can help researchers and practitioners in digital libraries to choose appropriate methods for automatic indexing.