Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research.

Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research.
复制标题

DOI:
10.1186/s12859-015-0472-9
复制
发表时间:
2015-02-21
期刊:
影响因子:
3
通讯作者:
Furlong LI
Furlong LI
中科院分区:
生物学4区
文献类型:
--
作者:
Bravo À;Piñero J;Queralt-Rosinach N;Rautschka M;Furlong LI

文献摘要

参考文献

被引文献

相似文献

当前的生物医学研究需要利用和利用科学出版物中报道的大量信息。自动文本挖掘方法,特别是那些旨在寻找实体之间的关系,是从自由文本库中识别可操作知识的关键。我们提出了BeFree系统,旨在确定生物医学实体之间的关系,特别关注基因及其相关疾病。通过利用文本的形态句法信息,BeFree能够以最先进的性能识别基因-疾病、药物-疾病和药物-靶点关联。BeFree在实际案例中的应用表明了其在翻译研究相关信息提取方面的有效性。我们通过大量分析和与其他数据源的整合,展示了BeFree提取的基因-疾病关联的价值。BeFree成功地识别了与全球发病率的主要原因抑郁症相关的基因,这些基因在其他公共资源中不存在。此外,基因-疾病关联的大规模提取和分析,以及与当前生物医学知识的整合,为文献中可以找到的信息类型提供了有趣的见解,并提出了有关数据优先级和策展的挑战。我们发现,只有一小部分通过使用BeFree发现的基因-疾病关联被收集在专家策划的数据库中。因此,迫切需要找到人工管理的替代策略,以审查,优先考虑和管理文本挖掘数据,并将其纳入特定领域的数据库。我们提出了我们的数据优先级的策略,并讨论其对支持生物医学研究和应用的影响。BeFree是一个新颖的文本挖掘系统,它在识别基因-疾病、药物-疾病和药物-靶点关联方面具有竞争力。我们的分析表明,仅挖掘MEDLINE的一小部分会产生一个大型的基因-疾病关联数据集,并且只有一小部分数据集实际上记录在策展资源中(2%),这引发了数据优先级和策展方面的几个问题。我们建议将文本挖掘数据与专家策划的数据进行联合分析,这似乎是评估数据质量并突出新颖有趣信息的合适方法。本文的在线版本(doi:10.1186/s12859-015-0472-9)包含补充材料,可供授权用户使用。
Current biomedical research needs to leverage and exploit the large amount of information reported in scientific publications. Automated text mining approaches, in particular those aimed at finding relationships between entities, are key for identification of actionable knowledge from free text repositories. We present the BeFree system aimed at identifying relationships between biomedical entities with a special focus on genes and their associated diseases. By exploiting morpho-syntactic information of the text, BeFree is able to identify gene-disease, drug-disease and drug-target associations with state-of-the-art performance. The application of BeFree to real-case scenarios shows its effectiveness in extracting information relevant for translational research. We show the value of the gene-disease associations extracted by BeFree through a number of analyses and integration with other data sources. BeFree succeeds in identifying genes associated to a major cause of morbidity worldwide, depression, which are not present in other public resources. Moreover, large-scale extraction and analysis of gene-disease associations, and integration with current biomedical knowledge, provided interesting insights on the kind of information that can be found in the literature, and raised challenges regarding data prioritization and curation. We found that only a small proportion of the gene-disease associations discovered by using BeFree is collected in expert-curated databases. Thus, there is a pressing need to find alternative strategies to manual curation, in order to review, prioritize and curate text-mining data and incorporate it into domain-specific databases. We present our strategy for data prioritization and discuss its implications for supporting biomedical research and applications. BeFree is a novel text mining system that performs competitively for the identification of gene-disease, drug-disease and drug-target associations. Our analyses show that mining only a small fraction of MEDLINE results in a large dataset of gene-disease associations, and only a small proportion of this dataset is actually recorded in curated resources (2%), raising several issues on data prioritization and curation. We propose that joint analysis of text mined data with data curated by experts appears as a suitable approach to both assess data quality and highlight novel and interesting information. The online version of this article (doi:10.1186/s12859-015-0472-9) contains supplementary material, which is available to authorized users.
用于蛋白质-蛋白质相互作用提取的步行加权子序列核。
DOI: 10.1186/1471-2105-11-107
发表时间: 2010-02-25
期刊: BMC bioinformatics
影响因子: 3
作者:
Kim S;Yoon J;Yang J;Park S
通讯作者: Park S
DOI: 10.1093/bioinformatics/btq538
发表时间: 2010-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bauer-Mehren, Anna;Rautschka, Michael;Furlong, Laura I.
通讯作者: Furlong, Laura I.
DOI: 10.1038/88213
发表时间: 2001-05-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Jenssen, TK;Lægreid, A;Hovig, E
通讯作者: Hovig, E
DOI: 10.1109/icimw.2009.5324773
发表时间: 2009-01-01
期刊: METHODS IN BIOENGINEERING: SYSTEMS ANALYSIS OF BIOLOGICAL NETWORKS
影响因子: --
作者:
Kim, Jin-Hong;Asthagiri, Anand R.
通讯作者: Asthagiri, Anand R.
DOI: 10.1093/bioinformatics/btm544
发表时间: 2008-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Kim, Seonho;Yoon, Juntae;Yang, Jihoon
通讯作者: Yang, Jihoon