Modern Public Health data storage for High Volume using the PowerVault MD3000
Modern Public Health data storage for High Volume using the PowerVault MD3000
批准号:
8052149
负责人:
Winston Alexander Hide
金额:
$47.31万
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-05-01 至 2013-04-30
关键词:
AreaBiological AssayBiologyBiometryCellsCommunicable DiseasesComplexComputational BiologyComputersCoupledDataData AnalysesData SetData Storage and RetrievalDecentralizationDisciplineDiseaseEnvironmentEnvironmental HealthEpidemiologyEtiologyFundingGenerationsGenesGeneticGenomeGenomicsGoalsHealthImmune responseIndividualInstitutionInvestigationModelingPerformancePhysical environmentPoliciesPopulationProcessPublic HealthPublic Health SchoolsResearchResourcesRoleScientistSocial EnvironmentSystemTechnologyUnited States National Institutes of Healthcohortcollaboratorycomputer infrastructurecostdata acquisitiondata managementdata sharingflexibilityimprovedmeetingsnutritionpathogenpublic health researchsocial
中文摘要
描述(由申请人提供):公共卫生研究越来越多地纳入高通量生物医学数据,为数据驱动的研究开辟了新的领域。最近,科学家们已经开始意识到现代生物学的潜力,即“超越基因组”,研究基因组与社会和物理环境的复杂相互作用,关注疾病病因学和所有细胞方面在促进健康方面的作用。为了实现这一潜力,我们的科学家们已经从个人的临时研究转向了旨在跨越广泛学科范围的合作项目。在过去五年中,环境和健康数据获取、测序和分析技术的规模急剧增加,加上数据产生工作日益分散,导致数据管理和分析的瓶颈日益增加。我们公共卫生学院的长期目标是提供一个无缝的合作环境,在这个环境中,我们可以利用从细胞到人群调查的共享数据集的广泛专业知识。为了实现这一目标,我们需要从根本上改进现有的共享计算机数据存储,从集中于低容量、高稳定性、高成本、高性能、用户支付所有成本的模式,到由机构补贴的分层数据存储模式,这种模式足够灵活,可以满足广泛的需求。我们希望:(a)在共享数据环境中共同定位基因组、遗传、环境、流行病学、社会和统计数据;(b)采用一致的政策、访问、用户支持、计算环境、工作流程和用户界面;(c)以低成本提供可扩展的数据存储资源,以适应基因组和队列数据规模的快速增长。因此,有效管理、存储和处理这些复杂的实验数据至关重要,需要能够提供一致存储和组织原始数据和衍生结果的计算基础设施。通过可扩展的共享数据存储,我们将直接影响复杂疾病的研究,宿主对传染病的反应,病原体多样性,营养以及基因对环境的研究。哈佛大学公共卫生学院(HSPH)正在申请资金,用于部署一个集中的、分层的高性能数据存储系统,以支持美国国立卫生研究院资助的计算生物学、基因组学和生物统计学应用于公共卫生的研究。
英文摘要
DESCRIPTION (provided by applicant): Public health research increasingly incorporates high-throughput biomedical data, opening up new areas for data-driven research. Recently, scientists have begun to realize the potential for modern biology to move 'beyond the genome' to look at the genome's complex interactions with the social and physical environments, focusing on disease etiology and the role of all cellular aspects in promoting health. In order to realize this potential our scientists have been moving from individual ad hoc studies to collaborative projects intended to scale across a broad range of disciplines. In the last five years, dramatic increases in the scale of environmental and health data acquisition, sequencing and assay technologies have coupled with increased decentralization of data generation resulting in a growing data management and analysis bottleneck. Our long term goal at the School of Public Health is to provide a seamless collaboratory environment in which it is possible to exploit the broad range of our expertise across shared datasets spanning investigations from the cell to the population. In order to achieve this aim we need to radically improve our existing shared computer data storage from its concentration on low volume, high stability, high cost, high performance with a user pays all costs model, to a tiered data storage model, subsidized by the institution, that is flexible enough to meet a broad range of requirements. We wish to: (a) co-locate genomic, genetic, environmental, epidemiological, social, and statistical data in a shared data environment; (b) apply consistent policies, access, user support, computing environments, workflows and user interfaces; ( c) provide a scalable data storage resource at low cost to accommodate the rapid increase in sizes of genomic and cohort data. The effective management, storage and processing of this complex experimental data is therefore crucial and requires computational infrastructure capable of providing consistent storage and organization of primary data and derived results. With scalable, shared data storage, we will directly impact studies in complex diseases, host response to infectious diseases, pathogen diversity, nutrition, and studies of genes to environment. The Harvard School of Public Health (HSPH) is requesting funding for the deployment of a centralized, tiered high-performance data storage system to support our NIH-funded research in computational biology, genomics and biostatistics as applied to public health.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Non-coding RNAs in resilience to Alzheimer’s Disease
-
批准号:10666167
-
项目类别:
-
资助金额:$86.25万
-
财政年份:2023
-
负责人:Winston Alexander Hide
-
依托单位:
The Alzheimer's Disease Resiliome: Pathway Analysis and Drug Discovery.
-
批准号:10374771
-
项目类别:
-
资助金额:$65.61万
-
财政年份:2019
-
负责人:Winston Alexander Hide
-
依托单位:
The Alzheimer's Disease Resiliome: Pathway Analysis and Drug Discovery.
-
批准号:10649411
-
项目类别:
-
资助金额:$65.38万
-
财政年份:2019
-
负责人:Winston Alexander Hide
-
依托单位:
海外基金