Textual Evidence Mining via Spherical Heterogeneous Information Network Embedding

Textual Evidence Mining via Spherical Heterogeneous Information Network Embedding
复制标题

DOI:
10.1109/bigdata50022.2020.9377958
复制
发表时间:
2020-12
期刊:
2020 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Xuan Wang;Yu Zhang;Aabhas Chauhan;Qi Li;Jiawei Han
Xuan Wang;Yu Zhang;Aabhas Chauhan;Qi Li;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Xuan Wang;Yu Zhang;Aabhas Chauhan;Qi Li;Jiawei Han

文献摘要

相似文献

科学文献作为主要知识资源之一,提供了丰富的文本证据,具有支持高质量科学假设验证的巨大潜力。在本文中,我们研究科学文献中的文本证据挖掘问题:给定一个科学假设作为查询三元组,在科学文献中找到支持输入查询的文本证据句子。科学文献中文本证据挖掘的一个关键挑战是在没有人类监督的情况下检索高质量的文本证据。因为获取大量包含科学文献中证据句子的人工注释文章并非易事。为了应对这一挑战,我们提出了 EvidenceMiner,这是一种用于科学文献的高质量文本证据检索方法,无需人工注释的训练示例。为了实现高质量的文本证据检索,我们利用现有知识库和大量非结构化文本中的异构信息。我们建议构建一个大型异构信息网络(HIN)来建立用户输入查询和候选证据句子之间的连接。基于构建的 HIN,我们提出了一种新颖的 HIN 嵌入方法,将节点直接嵌入到球形空间中以提高检索性能。对庞大的生物医学文献语料库(超过 400 万个句子)的定量实验表明,EvidenceMiner 显着优于无监督文本证据检索的基线方法。案例研究还表明,我们的 HIN 构建和嵌入极大地有利于许多下游应用,例如文本证据解释和同义词元模式发现。
Scientific literature, as one of the major knowledge resources, provides abundant textual evidence that has great potential to support high-quality scientific hypothesis validation. In this paper, we study the problem of textual evidence mining in scientific literature: given a scientific hypothesis as a query triplet, find the textual evidence sentences in scientific literature that support the input query. A critical challenge for textual evidence mining in scientific literature is to retrieve high-quality textual evidence without human supervision. Because it is non-trivial to obtain a large set of human-annotated articles containing evidence sentences in scientific literature. To tackle this challenge, we propose EvidenceMiner, a high-quality textual evidence retrieval method for scientific literature without human-annotated training examples. To achieve high-quality textual evidence retrieval, we leverage heterogeneous information from both existing knowledge bases and massive unstructured text. We propose to construct a large heterogeneous information network (HIN) to build connections between the user-input queries and the candidate evidence sentences. Based on the constructed HIN, we propose a novel HIN embedding method that directly embeds the nodes onto a spherical space to improve the retrieval performance. Quantitative experiments on a huge biomedical literature corpus (over 4 million sentences) demonstrate that EvidenceMiner significantly outperforms baseline methods for unsupervised textual evidence retrieval. Case studies also demonstrate that our HIN construction and embedding greatly benefit many downstream applications such as textual evidence interpretation and synonym meta-pattern discovery.