Automated cognome construction and semi-automated hypothesis generation

Automated cognome construction and semi-automated hypothesis generation
复制标题

DOI:
10.1016/j.jneumeth.2012.04.019
复制
发表时间:
2012-06-30
影响因子:
3
通讯作者:
Voytek, Bradley
Voytek, Bradley
中科院分区:
医学4区
文献类型:
--
作者:
Voytek, Jessica B.;Voytek, Bradley

文献摘要

被引文献

相似文献

现代神经科学研究站在无数巨人的肩膀上。仅PubMed就有2100多万篇同行评议的文章,每月有4 -5万篇以上的文章发表。了解人类的大脑、认知和疾病需要整合来自几十个科学领域的事实,这些事实分布在数百万个静态文件中的研究中,这让任何这样的整合都是令人望而生畏的。科学进步的未来将有助于弥合数百万篇已发表的研究论文与现代数据库(如Allen brain atlas (ABA))之间的差距。为此,我们分析了超过350万篇科学摘要的文本,以发现神经科学概念之间的联系。仅从文献中,我们就表明,我们可以盲目地通过算法提取“基因组”:大脑结构、功能和疾病之间的关系。通过引入两种半自动假设生成方法,我们展示了数据挖掘和与ABA跨平台数据集成的潜力。通过分析文献中的统计“漏洞”和差异,我们可以找到研究不足或被忽视的研究路径。也就是说,我们在科学过程本身的一部分上增加了一层半自动化。这是朝着从根本上将数据挖掘算法以一种可推广到任何科学或医学领域的方式纳入科学方法迈出的重要一步。(C) 2012 Elsevier B.V.版权所有
Modern neuroscientific research stands on the shoulders of countless giants. PubMed alone contains more than 21 million peer-reviewed articles with 40-50,000 more published every month. Understanding the human brain, cognition, and disease will require integrating facts from dozens of scientific fields spread amongst millions of studies locked away in static documents, making any such integration daunting, at best. The future of scientific progress will be aided by bridging the gap between the millions of published research articles and modern databases such as the Allen brain atlas (ABA). To that end, we have analyzed the text of over 3.5 million scientific abstracts to find associations between neuroscientific concepts. From the literature alone, we show that we can blindly and algorithmically extract a "cognome": relationships between brain structure, function, and disease. We demonstrate the potential of data-mining and cross-platform data-integration with the ABA by introducing two methods for semi-automated hypothesis generation. By analyzing statistical "holes" and discrepancies in the literature we can find understudied or overlooked research paths. That is, we have added a layer of semi-automation to a part of the scientific process itself. This is an important step toward fundamentally incorporating data-mining algorithms into the scientific method in a manner that is generalizable to any scientific or medical field. (C) 2012 Elsevier B.V. All rights reserved.