Big data from small data: data-sharing in the 'long tail' of neuroscience.

Big data from small data: data-sharing in the 'long tail' of neuroscience.
复制标题

DOI:
10.1038/nn.3838
复制
发表时间:
2014-11
影响因子:
25
通讯作者:
Martone ME
Martone ME
中科院分区:
医学1区
文献类型:
--
作者:
Ferguson AR;Nielson JL;Cragin MH;Bandrowski AE;Martone ME

文献摘要

被引文献

相似文献

美国大脑和欧洲人脑项目的启动恰逢国际社会日益努力提高透明度和增加获得公共资助的神经科学研究的机会。对数据共享标准和神经信息学基础设施的需求比以往任何时候都更加迫切。然而,“大科学”的努力并不是数据共享需求的唯一驱动力,因为神经科学家在整个研究领域都在努力应对每天产生的大量数据和越来越注重协作的科学环境。在这篇评论中,我们考虑了共享由单个神经科学家产生的丰富多样和异构的小数据集的问题,即所谓的长尾数据。我们考虑这些数据的实用性,存储库的多样性和可用于共享这些数据的选项,以及新兴的最佳实践。我们提供了一些用例,在这些用例中,聚合和挖掘不同的长尾数据将大量的小数据源转换为大数据,以提高对神经科学相关疾病的了解。
The launch of the US BRAIN and European Human Brain Projects coincides with growing international efforts toward transparency and increased access to publicly funded research in the neurosciences. The need for data-sharing standards and neuroinformatics infrastructure is more pressing than ever. However, ‘big science’ efforts are not the only drivers of data-sharing needs, as neuroscientists across the full spectrum of research grapple with the overwhelming volume of data being generated daily and a scientific environment that is increasingly focused on collaboration. In this commentary, we consider the issue of sharing of the richly diverse and heterogeneous small data sets produced by individual neuroscientists, so-called long-tail data. We consider the utility of these data, the diversity of repositories and options available for sharing such data, and emerging best practices. We provide use cases in which aggregating and mining diverse long-tail data convert numerous small data sources into big data for improved knowledge about neuroscience-related disorders.