DISEASES: Text mining and data integration of disease-gene associations

DISEASES: Text mining and data integration of disease-gene associations
复制标题

DOI:
10.1016/j.ymeth.2014.11.020
复制
发表时间:
2015-03-01
期刊:
影响因子:
4.8
通讯作者:
Jensen, Lars Juhl
Jensen, Lars Juhl
中科院分区:
生物学3区
文献类型:
--
作者:
Pletscher-Frankild, Sune;Palleja, Albert;Jensen, Lars Juhl

文献摘要

被引文献

相似文献

文本挖掘是一种灵活的技术,可以应用于生物学和医学中的许多不同任务。我们提出了一个从生物医学摘要中提取疾病-基因关联的系统。该系统由一个高效的基于字典的标注器组成,用于人类基因和疾病的命名实体识别,我们将其与一个考虑句子内部和句子之间共现的评分方案相结合。我们表明,这种方法能够提取一半的人工整理的关联,假阳性率仅为0.16%。尽管如此,文本挖掘不应该单独存在,而应该与其他类型的证据相结合。出于这个原因,我们开发了疾病资源,它将文本挖掘的结果与人工管理的疾病基因关联、癌症突变数据和现有数据库中的全基因组关联研究相结合。疾病资源可通过http://diseases.jensenlab.org/的web界面访问,其中文本挖掘软件和所有关联也可免费下载。(C) 2014年作者。Elsevier Inc.出版。
Text mining is a flexible technology that can be applied to numerous different tasks in biology and medicine. We present a system for extracting disease-gene associations from biomedical abstracts. The system consists of a highly efficient dictionary-based tagger for named entity recognition of human genes and diseases, which we combine with a scoring scheme that takes into account co-occurrences both within and between sentences. We show that this approach is able to extract half of all manually curated associations with a false positive rate of only 0.16%. Nonetheless, text mining should not stand alone, but be combined with other types of evidence. For this reason, we have developed the DISEASES resource, which integrates the results from text mining with manually curated disease-gene associations, cancer mutation data, and genome-wide association studies from existing databases. The DISEASES resource is accessible through a web interface at http://diseases.jensenlab.org/, where the text-mining software and all associations are also freely available for download. (C) 2014 The Authors. Published by Elsevier Inc.