DATA CORE
DATA CORE
批准号:
10686978
负责人:
JOSHUA T VOGELSTEIN
金额:
$23.47万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
未结题
起止时间:
2017-09-25 至 2027-08-31
关键词:
3-DimensionalAccelerationAddressAlgorithmsAnatomyAnimalsArchitectureAtlasesBehaviorBehavioralBig DataBiological ModelsBiologyBrainBrain imagingBrain regionCardiacCardiovascular systemCloud ComputingCodeCollaborationsCommunitiesComputer softwareDataData CollectionData Management ResourcesData ScienceData Science CoreData SetData SourcesDocumentationEcosystemElectron MicroscopyEnsureEnvironmentExperimental DesignsFAIR principlesFaceFeedbackGenerationsGoalsHeartImageIndividualInfrastructureIngestionInstitutionInternationalInternetKnowledgeLabelLearningLinkMapsMeasuresMemoryMetadataMindModalityModelingModernizationMorphologic artifactsNatureNeural InterconnectionNeuroanatomyNeurosciencesOntologyOxygenPathway interactionsPlantsProcessRegulationReproducibilityResearchResolutionResourcesRoleRunningScienceScientistServicesShapesSignal TransductionSliceSoftware EngineeringSurveysSystemTechnologyTestingTextureUnited States National Institutes of HealthVisualizationWritingZebrafishanalysis pipelinecollaboratorycommunity based researchcomparativecomputerized data processingcomputing resourcesconnectomedata curationdata integrationdata managementdata sharingdesigndigitalempowermentexperimental studygenomic dataimprovedinsightlaptopmembermigrationmind body interactionmodel buildingneural circuitnovelopen sourceopen source libraryportabilityrepositoryterabytetoolvirtual machineweb platform
中文摘要
数据科学核心-摘要
实现总体研究战略的科学目标需要在以下方面做出重大努力和取得进展:
神经科学的数据科学特别是,科学进步依赖于新颖的实验设计,数据
收集和处理(如项目1、2和3中所述),以及新的分析和模型(如
项目1、2和3),这导致要测试的一般原则(如项目2和3所述)。的
数据科学核心的基本目标是加速将所有收集的原始数据连接起来的过程。
三个项目用于获得数据衍生物的分析,然后可以用于构建跨
所有三个项目,并通过项目3中的电子显微镜(EM)进行验证。我们面临的两大挑战
加速这些联系的是大数据和可复制性。首先,收集的数据太大,无法装入内存,
或者甚至在磁盘上,每个实验的顺序为1 TB,整个数据集积累了数百个
TB或更多。因此,将MATLAB用于本地存储的所有分析的经典范例不是
足够了。解决这个问题的方法有两个:1)构建可扩展的算法,以便不同的人可以应用它们
这些大数据,2)开发云数据管理系统,以便所有联盟成员能够快速
访问和分析数据,然后将它们彼此集成。云数据管理系统
将建立在为Open Connectome项目1开发的基础设施上,该项目最初是为托管数据而开发的
以及ZBrain 2.0,我们正在开发的一种资源,用于定义一个共同的坐标空间
用于斑马鱼大脑图谱绘制。其次,这是一个团队的努力,所以分享分析和衍生品,并保持跟踪
元数据将非常重要。解决这一问题的办法有三:1)建立一个全面的科学环境
在云中,这使得整个“数字实验”的共享,链接到数据,并确保整个
分析管道可以由任何人和任何地方进行日常运行和扩展,2)仔细管理数据,
现有资源中的元数据,以及3)促进不同成像数据集的集成以改进ZBrain
2.0.我们的整个系统是建立在并将继续是开源的,可移植的和可复制的,并将使用
数据科学和FAIR的最佳实践(
可查找、可扩展、可互操作和可重用)2
数据
管理完成这个数据科学核心的所有目标不仅可以实现和加速科学发展,
通过这项提案所解决的进展,它将建立数据科学的新标准,
适用于所有其他U19的努力,以及许多其他努力内外NIH,甚至国际
科学的努力在大。
英文摘要
Data Science Core - Abstract
Achieving the scientific goals of the Overall Research Strategy requires a significant effort and advancement in
data science for neuroscience. In particular, scientific progress depends on novel experimental design, data
collection and processing (as described in Projects 1, 2, and 3), and novel analysis and models (as described in
Projects 1, 2, and 3), which lead to general principles to be tested (as described in Projects 2 and 3). The
fundamental goal of the Data Science Core is to accelerate the process connecting the raw data collected in all
three Projects to the analyses used to obtain data derivatives, which can then be used to build models across
all three Projects, and validated via electron microscopy (EM) in Project 3. The two main challenges we face to
accelerate these links are big data and reproducibility. First, the data collected are too large to fit into memory,
or even on disk, with each experiment ordering on one terabyte (TB), and the entire dataset amassing hundreds
of TB or more. Therefore, the classic paradigm of using MATLAB for all analyses that are stored locally is not
sufficient. The solution to this is twofold: 1) build scalable algorithms, so that different individuals can apply them
to these big data, and 2) develop cloud data management systems, so that all consortium members can quickly
access and analyze the data, and then integrate them with one another. The cloud data management system
will be built on the infrastructure developed for the Open Connectome Project1, originally developed to host data
on institutional resources, and ZBrain 2.0, a resource we are developing to define a common coordinate space
for zebrafish brain atlasing. Second, this is a team effort, so sharing analyses and derivatives and keeping track
of metadata will be important. The solution to this is threefold: 1) build a comprehensive scientific environment
in the cloud, that enables sharing of entire “digital experiments”, linking to the data and ensuring that the entire
analysis pipeline can be trivially run and extended by anyone and anywhere, 2) carefully curating data and
metadata in existing resources, and 3) facilitating the integration of different imaging datasets to improve ZBrain
2.0. Our entire system is built on and will continue to be open source, portable and reproducible, and will use
and extend best practices of data science and FAIR (
Findable, Accessible, Interoperable, and Re-usable)2
data
management. Completing all the aims in this Data Science Core will not only enable and accelerate the scientific
progress addressed by this proposal, it will establish new standards in data science that can be immediately
applied to all other U19 efforts, as well as many other efforts within and outside NIH and even the international
science effort at large.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
DATA CORE
-
批准号:10525429
-
项目类别:
-
资助金额:$26.89万
-
财政年份:2017
-
负责人:JOSHUA T VOGELSTEIN
-
依托单位:
Data Science Core
-
批准号:10241479
-
项目类别:
-
资助金额:$45.29万
-
财政年份:2017
-
负责人:JOSHUA T VOGELSTEIN
-
依托单位:
Data Science Core
-
批准号:9444234
-
项目类别:
-
资助金额:$34.59万
-
财政年份:--
-
负责人:JOSHUA T VOGELSTEIN
-
依托单位:
海外基金