Large-scale extraction of gene interactions from full-text literature using DeepDive.

Large-scale extraction of gene interactions from full-text literature using DeepDive.
复制标题

DOI:
10.1093/bioinformatics/btv476
复制
发表时间:
2016-01-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Altman RB
Altman RB
中科院分区:
其他
文献类型:
--
作者:
Mallory EK;Zhang C;Ré C;Altman RB

文献摘要

被引文献

相似文献

动机:基因间相互作用的完整存储库是理解细胞过程、人类疾病和药物反应的关键。这些基因-基因相互作用包括蛋白质-蛋白质相互作用和转录因子相互作用。大多数已知的相互作用都可以在生物医学文献中找到。相互作用数据库,例如 BioGRID 和 ChEA,注释了这些基因-基因相互作用;然而,随着文献呈指数级增长,策展变得困难。 DeepDive 是一个经过训练的系统,用于从各种来源(包括文本)提取信息。在这项工作中,我们使用 DeepDive 从超过 100000 篇全文 PLOS 文章中提取蛋白质-蛋白质和转录因子相互作用。方法:我们构建了一个基因-基因相互作用提取器,用于识别输入句子中的候选基因-基因关系。对于每个候选关系,DeepDive 计算该关系是正确交互的概率。我们根据相互作用蛋白质数据库和随机提取的数据对该系统进行了评估。结果:我们的系统在提取涉及句子中同时出现的基因符号的直接​​和间接相互作用时,准确率达到 76%,召回率达到 49%。对于随机策划的提取,系统基于直接或间接交互以及句子级和文档级精度实现了 62% 到 83% 的精度。总体而言,我们的系统使用 100 万多篇全文文章中的 724 个特征提取了 3356 个独特的基因对。可用性和实施​​:应用程序源代码可在 https://github.com/edoughty/deepdive_genegene_app 上公开获取。 联系方式:russ.altman@stanford.edu 补充信息:补充数据可在 Bioinformatics 在线获取。
Motivation: A complete repository of gene–gene interactions is key for understanding cellular processes, human disease and drug response. These gene–gene interactions include both protein–protein interactions and transcription factor interactions. The majority of known interactions are found in the biomedical literature. Interaction databases, such as BioGRID and ChEA, annotate these gene–gene interactions; however, curation becomes difficult as the literature grows exponentially. DeepDive is a trained system for extracting information from a variety of sources, including text. In this work, we used DeepDive to extract both protein–protein and transcription factor interactions from over 100 000 full-text PLOS articles. Methods: We built an extractor for gene–gene interactions that identified candidate gene–gene relations within an input sentence. For each candidate relation, DeepDive computed a probability that the relation was a correct interaction. We evaluated this system against the Database of Interacting Proteins and against randomly curated extractions. Results: Our system achieved 76% precision and 49% recall in extracting direct and indirect interactions involving gene symbols co-occurring in a sentence. For randomly curated extractions, the system achieved between 62% and 83% precision based on direct or indirect interactions, as well as sentence-level and document-level precision. Overall, our system extracted 3356 unique gene pairs using 724 features from over 100 000 full-text articles. Availability and implementation: Application source code is publicly available at https://github.com/edoughty/deepdive_genegene_app Contact: russ.altman@stanford.edu Supplementary information: Supplementary data are available at Bioinformatics online.