课题基金 / 基金详情

项目摘要

项目成果

JOSHUA T VOGELSTEIN的其他基金

相似基金

相关文献

中文摘要
翻译
数据科学核心-摘要 实现总体研究战略的科学目标需要在数据方面做出重大努力和取得进展 神经科学的科学。特别是,科学进步取决于新颖的实验设计、数据收集和 处理(如项目1和2所述)以及新的分析和模型(如项目3所述),这导致 要测试的一般原则(如项目4中所述)。数据科学核心的根本目标是加速 将项目1、2和4中收集的原始数据与用于获得数据派生数据的分析联系起来的过程, 然后可以用来在项目3中构建模型,并在项目4中扩展它们。 这些链接是大数据和可再现性。首先,收集的数据太大,无法放入内存,甚至无法存储在磁盘上 每个实验都以1TB为单位进行订购,整个数据集累积到数百TB或更多。因此, 使用MatLab进行本地存储的所有分析的经典范例是不够的。解决这个问题的办法是 双重:(1)构建云数据管理系统,让所有联合体成员能够快速访问和分析 数据,以及(2)构建可扩展的算法,以便不同的个人可以将它们应用于这些大数据。云数据 管理系统将建立在为最初开发的Open Connectome Project 1​​开发的基础设施上 托管有关机构资源的数据。在过去的一年里,团队已经成熟,成为了NeuroData(http://neurodata.io​​), 将所有基础架构移植到商业云,并已托管20个数据集,共50 TB,包括所有 这里提出了三种分析尺度(h​ttp://urodata.io​)。可伸缩的算法将基于来自 Http://flashx.io称其为FlashX(​NeuroData​)。FlashX是一个C图形分析和机器学习库,旨在运行 仅使用一台计算机(不是群集)、2台​、3台​和最近收到的DARPA对任意大数据进行分析 SBIR奖商业化。我们将使用FlashX作为后端,以支持所有处理行为和 成像数据。其次,这是团队的努力,因此共享分析和衍生工具以及跟踪元数据将是 很重要。解决方案是在云中构建一个全面的科学环境,以实现共享 整个“数字实验”,链接到数据,并确保整个分析管道可以平凡地运行和 任何人和任何地方都可以扩展。该系统将扩展NeuroData的“云中科学” (http://scienceinthe.cloud​​)4​,5​,最近接受了私人资金的专业化。我们的整个系统都建立在 将继续是开源的、可移植的和可重现的,并将使用和推广数据科学和 公平(​​数据管理。完成这门数据科学的所有目标 可查找、可访问、可互操作和可重复使用) CORE不仅将推动和加速这项提议所涉及的科学进步,它还将建立新的标准 在数据科学方面,可以立即应用于所有其他U19努力,就像NIH内外的许多其他努力一样 甚至整个国际科学努力也是如此。
英文摘要
Data Science Core-Abstract Achieving the scientific goals of the Overall Research Strategy requires a significant effort and advancement in data science for neuroscience. In particular, scientific progress depends on novel experimental design, data collection and processing (as described in Projects 1 and 2), and novel analysis and models (as described in Project 3), which lead to general principles to be tested (as described in Project 4). The fundamental goal of the Data Science Core is to accelerate the process connecting the raw data collected in Projects 1, 2, and 4 to the analyses used to obtain data derivatives, which can then be used to build models in Project 3, and extend them in Project 4. The two main challenges we face to accelerate these links are big data and reproducibility. First, the data collected are too large to fit into memory, or even on disk, with each experiment ordering on one terabyte (TB), and the entire dataset amassing hundreds of TB or more. Therefore, the classic paradigm of using MATLAB for all analyses that are stored locally is not sufficient. The solution to this is twofold: (1) build a cloud data management system, so that all consortium members can quickly access and analyze the data, and (2) build scalable algorithms, so that different individuals can apply them to these big data. The cloud data management system will be built on the infrastructure developed for the Open Connectome Project 1​ ​, originally developed to host data on institutional resources. In the last year, the team has matured to become NeuroData (​http://neurodata.io​), porting all the infrastructure to the commercial cloud, and already hosting 20+ datasets comprising 50+ TB, including all three scales of analysis proposed here (h​ ttp://neurodata.io​). The scalable algorithms will be based on another project from NeuroData called FlashX (​http://flashx.io​). FlashX is a C++ graph analytics and machine learning library, designed to run analytics on arbitrarily large data using only a single machine (not a cluster) 2​ ,3​, and the recent recipient of a DARPA SBIR award to commercialize. We will use FlashX as a backend to support all the algorithms for processing behavior and imaging data. Second, this is a team effort, so sharing analyses and derivatives and keeping track of metadata will be important. The solution to this is to build a comprehensive scientific environment in the cloud, that enables sharing of entire “digital experiments”, linking to the data and ensuring that the entire analysis pipeline can be trivially run and extended by anyone and anywhere. This system will extend NeuroData’s “Science in the Cloud” (​http://scienceinthe.cloud​) 4​ ,5​, which recently received private funding to professionalize. Our entire system is built on and will continue to be open source, portable and reproducible, and will use and extend best practices of data science and FAIR (​ ​data management. Completing all the aims in this Data Science Findable, Accessible, Interoperable, and Re-usable) Core will not only enable and accelerate the scientific progress addressed by this proposal, it will establish new standards in data science that can be immediately applied to all other U19 efforts, as many other efforts within and outside NIH and even the international science effort at large.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
DATA CORE
  • 批准号:
    10525429
  • 项目类别:
  • 资助金额:
    $26.89万
  • 财政年份:
    2017
  • 负责人:
    JOSHUA T VOGELSTEIN
  • 依托单位:
DATA CORE
  • 批准号:
    10686978
  • 项目类别:
  • 资助金额:
    $23.47万
  • 财政年份:
    2017
  • 负责人:
    JOSHUA T VOGELSTEIN
  • 依托单位:
Data Science Core
  • 批准号:
    10241479
  • 项目类别:
  • 资助金额:
    $45.29万
  • 财政年份:
    2017
  • 负责人:
    JOSHUA T VOGELSTEIN
  • 依托单位:
海外基金