Content-rich biological network constructed by mining PubMed abstracts.

Content-rich biological network constructed by mining PubMed abstracts.
复制标题

DOI:
10.1186/1471-2105-5-147
复制
发表时间:
2004-10-08
期刊:
影响因子:
3
通讯作者:
Sharp BM
Sharp BM
中科院分区:
生物学4区
文献类型:
--
作者:
Chen H;Sharp BM

文献摘要

参考文献

被引文献

相似文献

由强大的技术进步(例如微阵列)以及来自多个物种的基因组序列的可用性而产生的有关基因组,转录组和蛋白质组的迅速扩展的信息语料库的整合,挑战了科学界的掌握和理解。尽管存在基于基因/蛋白质术语或摘要文本中相似性的文本共发生识别生物学关系的文本挖掘方法,但大规模了解潜在的分子连接的知识,这是理解新型生物学过程的先决条件,落后于数据的积累。虽然在计算上有效,但基于共发生的方法无法表征(例如抑制或刺激,方向性)生物学相互作用。已经创建了具有自然语言处理(NLP)功能的程序来解决这些限制,但是,它们通常不容易被公众访问。 我们提出了一种基于NLP的文本挖掘方法Chilibot,该方法在生物学概念,基因,蛋白质或药物之间构建了富含内容的关系网络。在其特征中,可以提出有关新假设的建议。最后,我们提供的证据表明,从生物学文献中提取的分子网络的连通性遵循幂律分布,表明与先前实验分析的结果一致。 Chilibot从各种生物领域的知识中提取科学关系,并以内容丰富的图形格式呈现它们,从而将一般的生物医学知识与用户的专业知识和兴趣相结合。可以向学术用户免费访问Chilibot。
The integration of the rapidly expanding corpus of information about the genome, transcriptome, and proteome, engendered by powerful technological advances, such as microarrays, and the availability of genomic sequence from multiple species, challenges the grasp and comprehension of the scientific community. Despite the existence of text-mining methods that identify biological relationships based on the textual co-occurrence of gene/protein terms or similarities in abstract texts, knowledge of the underlying molecular connections on a large scale, which is prerequisite to understanding novel biological processes, lags far behind the accumulation of data. While computationally efficient, the co-occurrence-based approaches fail to characterize (e.g., inhibition or stimulation, directionality) biological interactions. Programs with natural language processing (NLP) capability have been created to address these limitations, however, they are in general not readily accessible to the public. We present a NLP-based text-mining approach, Chilibot, which constructs content-rich relationship networks among biological concepts, genes, proteins, or drugs. Amongst its features, suggestions for new hypotheses can be generated. Lastly, we provide evidence that the connectivity of molecular networks extracted from the biological literature follows the power-law distribution, indicating scale-free topologies consistent with the results of previous experimental analyses. Chilibot distills scientific relationships from knowledge available throughout a wide range of biological domains and presents these in a content-rich graphical format, thus integrating general biomedical knowledge with the specialized knowledge and interests of the user. Chilibot can be accessed free of charge to academic users.
DOI: 10.1093/bioinformatics/btg207
发表时间: 2003-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Novichkova, S;Egorov, S;Daraselia, N
通讯作者: Daraselia, N
DOI: 10.1038/88213
发表时间: 2001-05-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Jenssen, TK;Lægreid, A;Hovig, E
通讯作者: Hovig, E
DOI: 10.1038/35075138
发表时间: 2001-05-03
期刊: NATURE
影响因子: 64.8
作者:
Jeong, H;Mason, SP;Oltvai, ZN
通讯作者: Oltvai, ZN
DOI: 10.1093/bioinformatics/17.10.988
发表时间: 2001-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rzhetsky, A;Gomez, SM
通讯作者: Gomez, SM
DOI: 10.1038/35036627
发表时间: 2000-10-05
期刊: NATURE
影响因子: 64.8
作者:
Jeong, H;Tombor, B;Barabási, AL
通讯作者: Barabási, AL