Demystifying the Semantics of Relevant Objects in Scholarly Collections: A Probabilistic Approach

Demystifying the Semantics of Relevant Objects in Scholarly Collections: A Probabilistic Approach
复制标题

DOI:
10.1145/2756406.2756923
复制
发表时间:
2015-06
期刊:
Proceedings of the 15th ACM/IEEE-CS Joint Conference on Digital Libraries
影响因子:
--
通讯作者:
J. M. Pinto;Wolf-Tilo Balke
J. M. Pinto;Wolf-Tilo Balke
中科院分区:
其他
文献类型:
--
作者:
J. M. Pinto;Wolf-Tilo Balke

文献摘要

相似文献

通过科学数字图书馆获取高度专业知识的努力需要超越仅仅的书目元数据,因为这里的信息搜索大多以实体为中心。以前的工作已经认识到这一趋势,并开发了不同的方法来识别和(在某种程度上甚至自动)注释几种重要类型的实体:基因和蛋白质,化学结构和分子,或药物名称,仅举几例。此外,这些实体经常与精选数据库中的条目交叉引用。然而,有几个问题仍然有待回答:给定一个科学学科,什么是重要的实体?如何自动识别它们?它们真的都是相关的吗?也就是说,它们都带有评估出版物的更深层次的语义吗?它们如何被表示、描述和随后的注释?如何将它们用于搜索任务?在这项工作中,我们专注于回答其中的一些问题。我们主张,要将科学数字图书馆的使用提升到一个新的水平,我们必须找到将特定主题实体视为一等公民,并将其语义深度整合到搜索过程中。为了支持这一点,我们提出了一种新的概率方法,不仅成功地提供了一个解决方案的集成问题,而且还演示了如何利用编码在实体中的知识,并提供见解,探索我们的方法在不同的情况下使用。最后,我们展示了我们的结果如何使信息提供者受益。
Efforts to make highly specialized knowledge accessible through scientific digital libraries need to go beyond mere bibliographic metadata, since here information search is mostly entity-centric. Previous work has realized this trend and developed different methods to recognize and (to some degree even automatically) annotate several important types of entities: genes and proteins, chemical structures and molecules, or drug names to name but a few. Moreover, such entities are often crossreferenced with entries in curated databases. However, several questions still remain to be answered: Given a scientific discipline what are the important entities? How can they be automatically identified? Are really all of them relevant, i.e. do all of them carry deeper semantics for assessing a publication? How can they be represented, described, and subsequently annotated? How can they be used for search tasks? In this work we focus on answering some of these questions. We claim that to bring the use of scientific digital libraries to the next level we must find treat topic-specific entities as first class citizens and deeply integrate their semantics into the search process. To support this we propose a novel probabilistic approach that not only successfully provides a solution to the integration problem, but also demonstrates how to leverage the knowledge encoded in entities and provide insights to explore the use of our approach in different scenarios. Finally, we show how our results can benefit information providers.