Analysis of Biological Processes and Diseases Using Text Mining Approaches

Analysis of Biological Processes and Diseases Using Text Mining Approaches
复制标题

DOI:
10.1007/978-1-60327-194-3_16
复制
发表时间:
2010-01-01
期刊:
BIOINFORMATICS METHODS IN CLINICAL RESEARCH
影响因子:
--
通讯作者:
Valencia, Alfonso
Valencia, Alfonso
中科院分区:
其他
文献类型:
--
作者:
Krallinger, Martin;Leitner, Florian;Valencia, Alfonso

文献摘要

被引文献

相似文献

已经开发了许多生物医学文本挖掘系统来直接从文献中提取生物学相关信息,补充了分析实验生成的数据中的生物信息学方法。我们简要概述了自然语言数据的一般特征、现有的生物医学文献数据库以及生物医学文本挖掘背景下的相关词汇资源。介绍了选定数量的实际有用的系统以及支持的用户查询类型及其生成的结果。将通过癌症相关蛋白质的示例案例来讨论使用信息提取系统提取生物关系,例如蛋白质-蛋白质相互作用以及代谢和信号传导途径。描述了检测基因与疾病关联的基本策略以及突变、SNP 和表观遗传信息(甲基化)的文献挖掘。我们概述了以疾病为中心和以基因为中心的文献挖掘方法,用于将基因与表型和基因型方面联系起来。此外,我们还讨论了最近通过文本挖掘寻找生物标志物以及基因列表分析和优先级排序的努力。将指出实现定制生物医学文本挖掘系统的索尼克相关问题。为了证明文献挖掘对于分子肿瘤学领域的有用性,我们实施了两个与癌症相关的应用程序。第一个工具由一个文献挖掘系统组成,用于检索人类突变以及支持文章。特定的基因突变与一组预先定义的癌症类型相关。第二个应用程序包括一个文本分类系统,支持乳腺癌特定的文献搜索和基于文档的乳腺癌基因排名。文本挖掘的未来趋势强调社区努力的重要性,例如 BioCreative 挑战,用于将多个系统开发和集成到 BioCreative Metaserver 提供的通用平台中。
A number of biomedical text mining systems have been developed to extract biologically relevant information directly from the literature, complementing bioinformatics methods in the analysis of experimentally generated data. We provide a short overview of the general characteristics of natural language data, existing biomedical literature databases, and lexical resources relevant in the context of biomedical text mining. A selected number of practically useful systems are introduced together with the type of user queries Supported and the results they generate. The extraction of biological relationships, such as protein-protein interactions as well as metabolic and signaling pathways using information extraction systems, will be discussed through example cases of cancer-relevant proteins. Basic strategies for detecting associations of genes to diseases together with literature mining of mutations, SNPs, and epigenetic information (methylation) are described. We provide an overview of disease-centric and gene-centric literature mining methods for linking genes to phenotypic and genotypic aspects. Moreover, we discuss recent efforts for finding biomarkers through text mining and for gene list analysis and prioritization. Sonic relevant issues for implementing a customized biomedical text mining system will be pointed out. To demonstrate the usefulness of literature mining for the molecular oncology domain, we implemented two cancer-related applications. The first tool consists of a literature mining system for retrieving human mutations together with Supporting articles. Specific gene mutations are linked to a set of predefined cancer types. The second application consists of a text categorization system supporting breast cancer-specific literature search and document-based breast cancer gene ranking. Future trends in text mining emphasize the importance of community efforts such as the BioCreative challenge for the development and integration Of multiple Systems into a common platform provided by the BioCreative Metaserver.