Automatic Human-like Mining and Constructing Reliable Genetic Association Database with Deep Reinforcement Learning

Automatic Human-like Mining and Constructing Reliable Genetic Association Database with Deep Reinforcement Learning
复制标题

DOI:
10.1101/434803
复制
发表时间:
2018-10
影响因子:
--
通讯作者:
Haohan Wang;Xiang Liu;Yifeng Tao;Wenting Ye;Qiao Jin;William W. Cohen;E. Xing
Haohan Wang;Xiang Liu;Yifeng Tao;Wenting Ye;Qiao Jin;William W. Cohen;E. Xing
中科院分区:
--
文献类型:
--
作者:
Haohan Wang;Xiang Liu;Yifeng Tao;Wenting Ye;Qiao Jin;William W. Cohen;E. Xing

文献摘要

相似文献

生物和生物医学科学研究中越来越多的科学文献对持续可靠地管理最新发现的知识提出了挑战,而自动生物医学文本挖掘已经成为应对这一挑战的答案之一。在本文中,我们的目标是通过训练系统直接模拟人类的行为,如查询PubMed,从查询结果中选择文章,以及阅读所选文章以获取知识,从而进一步提高生物医学文本挖掘的可靠性。我们利用生物医学文本挖掘的效率,深度强化学习的灵活性,以及UMLS中收集的大量知识,将其集成为一个人工智能阅读器,可以自动识别真实的文章,并有效地获取文章中所传达的知识。我们构建了一个系统,其当前的主要任务是建立基因与人类复杂性状之间的遗传关联数据库。我们在本文中的贡献有三个方面:1)我们提出通过构建一个可以直接模拟研究人员行为的系统来提高文本挖掘的可靠性,我们开发了相应的方法,如用于文本挖掘的双向LSTM和用于组织行为的Deep Q-Network。2)以构建遗传关联数据库为例,验证了该系统的有效性。3)我们将我们的实现作为一个通用框架发布,供社区的研究人员方便地构建其他数据库。
The increasing amount of scientific literature in biological and biomedical science research has created a challenge in the continuous and reliable curation of the latest knowledge discovered, and automatic biomedical text-mining has been one of the answers to this chal-lenge. In this paper, we aim to further improve the reliability of biomedical text-mining by training the system to directly simulate the human behaviors such as querying the PubMed, selecting articles from queried results, and reading selected articles for knowledge. We take advantage of the efficiency of biomedical text-mining, the flexibility of deep reinforcement learning, and the massive amount of knowledge collected in UMLS into an integrative arti-ficial intelligent reader that can automatically identify the authentic articles and effectively acquire the knowledge conveyed in the articles. We construct a system, whose current pri-mary task is to build the genetic association database between genes and complex traits of the human. Our contributions in this paper are three-fold: 1) We propose to improve the reliability of text-mining by building a system that can directly simulate the behavior of a researcher, and we develop corresponding methods, such as Bi-directional LSTM for text mining and Deep Q-Network for organizing behaviors. 2) We demonstrate the effec-tiveness of our system with an example in constructing a genetic association database. 3) We release our implementation as a generic framework for researchers in the community to conveniently construct other databases.