Investigation of Cloud Computing to Support Data-Parallel Health Research
Investigation of Cloud Computing to Support Data-Parallel Health Research
批准号:
7944073
负责人:
GEOFFREY C FOX
金额:
$75.63万
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-30 至 2011-08-31
关键词:
AddressAlgorithmsApache IndiansArchitectureAreaBase SequenceBiologicalBiological SciencesBiologyCerealsCollaborationsCommunitiesComputational BiologyComputersDNA Sequence AnalysisDataData AnalysesData SetData Storage and RetrievalDevicesFaceFundingFunding OpportunitiesGene Expression ProfileGenesGenomeGenomicsGoalsHealthImageryIndianaInformation RetrievalInternetInvestigationMapsMetagenomicsMutationOccupationsOutsourcingPatientsPerformancePopulationProcessProductionReadingResearchResearch InfrastructureResearch Project GrantsResourcesRunningScientistServicesSourceSystemTechniquesTechnologyUnited States National Institutes of HealthUniversitiesWorkbasecluster computingcomparative genomicscomputerized data processingdata mininghealth recordmultidisciplinaryparallel processingprogramspublic health relevancesupercomputertool
中文摘要
描述(由申请人提供):我们将组建一个由印第安纳州大学计算机科学家、生物学家和生物信息学家组成的多学科团队,开发和部署新的大规模计算基础设施和工具,以实现基础健康研究。我们的研究将调查云计算架构对大规模计算生物学的影响,特别是广泛遇到的“数据并行”问题,包括但不限于DNA序列分析。GO基金将用于建立基于云计算的生命科学新领域。云计算目前以Amazon Web Services、Microsoft Azure和其他商业努力为代表。然而,许多大学(包括印第安纳州大学)正在建立研究云部署,这将解决两个一般性问题:基础设施:云提供简单的Web服务编程接口,允许科学家创建计算集群和使用高度可靠的数据存储。也就是说,云提供了一种外包计算基础设施的方式。 运行时:云系统特别适合运行大规模的信息检索问题。这些数据并行问题涉及到处理分成许多块的非常大的数据集的复制、顺序命令的管道。示例技术包括Microsoft Dryad和Apache Hadoop。在这项提案中,我们将与微软研究院合作,微软研究院目前正在将Dryad从一个研究项目转变为一个强大的工具。我们已经分析了各种各样的健康研究问题,并表明它们可以从云基础设施和运行时中受益。云为研究小组提供了一种外包计算、存储和网络的方式,并在健康研究中的数据并行问题上实现高性能。我们团队的研究成果(许多NIH资助)代表了广泛的应用,包括a)基于序列的转录组分析,B)基因组重测序突变图谱,c)宏基因组学分析,d)基因组注释,e)比较基因组学,f)群体基因组学h)患者健康记录中的高级并行数据挖掘。处理大规模数据是这些工作的共同问题。
公共卫生相关性: 我们建议调查和开发一个独特的云计算研究基础设施,这将对几个不同的生命科学研究领域产生非常大的影响。我们的重点是大规模的,数据并行分析的问题,导致从短读基因测序设备和其他来源的数据泛滥。我们将与现有的几个生物和生物医学项目合作开发和展示我们的基础设施。
英文摘要
DESCRIPTION (provided by applicant): We will form a multidisciplinary team of Indiana University computer scientists, biologists, and bioinformaticians to develop and deploy new large-scale computing infrastructure and tools that will enable fundamental health research. Our research will investigate the impact of Cloud computing architectures on large-scale computational biology, particularly widely encountered, "data parallel" problems including but not limited to DNA sequence analysis. GO funds will be used to establish the new field of Cloud-based computational life science. Cloud computing is currently typified by Amazon Web Services, Microsoft Azure, and other commercial efforts. However, many universities (including Indiana University) are in the process of establishing research Cloud deployments that will address two general problems: Infrastructure: Clouds provide simple Web service programming interfaces that allows scientists to create computing clusters and use highly reliable data storage. That is, Clouds provide a way to outsource computing infrastructure. Runtimes: Cloud systems are particularly appropriate for running large-scale information retrieval problems. These data-parallel problems involve pipelines of replicated, sequential commands that process very large data sets divided into many pieces. Example technologies include Microsoft Dryad and Apache Hadoop. In this proposal, we will partner with Microsoft Research, which is currently converting Dryad from a research project to a robust tool. We have analyzed a wide variety of health research problems and have shown that they can benefit from Cloud infrastructure and runtimes. Clouds provide research groups with a way to outsource computing, storage, and networking and to achieve high performance on data-parallel problems in health research. Our team's research efforts (many NIH funded) represent a wide range of applications, including a) sequence-based transcriptome profiling, b) genome re-sequencing for mutation mapping, c) metagenomics analysis, d) genome annotation, e) comparative genomics, and f) population genomics h) advanced parallel datamining in patient health records. Processing large-scale data is the common problem uniting these efforts.
PUBLIC HEALTH RELEVANCE: We propose to investigate and develop a unique Cloud computing research infrastructure that will have a very large impact on several different life science research areas. Our focus is on the large-scale, data-parallel analysis problems that result from the deluge of data from short-read gene sequencing devices and other sources. We will develop and demonstrate our infrastructure in collaboration with several existing biological and biomedical projects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Investigation of Cloud Computing to Support Data-Parallel Health Research
-
批准号:7852166
-
项目类别:
-
资助金额:$73.5万
-
财政年份:2009
-
负责人:GEOFFREY C FOX
-
依托单位:
Chemical Informatics Cyberinfrastructure (RMI)
-
批准号:7032188
-
项目类别:
-
资助金额:$36.33万
-
财政年份:2005
-
负责人:GEOFFREY C FOX
-
依托单位:
Chemical Informatics Cyberinfrastructure
-
批准号:7476645
-
项目类别:
-
资助金额:$34.28万
-
财政年份:2005
-
负责人:GEOFFREY C FOX
-
依托单位:
Chemical Informatics Cyberinfrastructure
-
批准号:7125590
-
项目类别:
-
资助金额:$36.85万
-
财政年份:2005
-
负责人:GEOFFREY C FOX
-
依托单位:
海外基金