课题基金 / 基金详情

项目摘要

项目成果

ANTON NEKRUTENKO的其他基金

相似基金

相关文献

中文摘要
翻译
“下一代”测序(NGS)仪器的广泛可用性使任何研究人员都能够以适度的成本产生大量的DNA序列数据。然而,使用这些原始序列对个体研究者、小型实验室或核心设施提出了重大问题。对于一个没有计算专业知识的实验组来说,简单地运行数据分析程序就是一个障碍,更不用说构建能够处理NGS数据的计算和数据存储基础设施了。幸运的是,最近出现了一种计算模型--“云计算”,非常适合于大规模序列数据的分析。在该模型中,计算和存储作为虚拟资源存在,可以根据需要动态分配和释放。重要的是,对于某些用例,云资源可以以比专用资源低得多的成本提供存储和计算。然而,要使调查人员能够利用这些资源,还需要克服巨大的挑战。具体来说,虽然云计算提供了一种按需获取计算资源的方式,但提供的资源要么是互联网上的虚拟机,要么是特定的编程库,实验者无法使用。因此,一个可行的分析解决方案需要在没有信息学专业知识的情况下也可以访问和部署;它必须有效和自动地使用动态可扩展的资源,同时考虑到时间和成本;它必须包括适当的分析工具,并在新工具出现时轻松支持添加。我们以前开发了一个软件系统-银河(http:galaxyproject.org)-提供了一个强大的框架,以满足这些需求。在这里,我们建议显着扩展这个框架,允许任何实验者利用云计算基础设施的力量进行大规模的NGS分析。特别是,我们将修改现有的Galaxy框架,使其完全在云中运行。我们将调整Galaxy调度和执行作业的方式,以有效利用云风格。我们将为个人用户提供一种机制,通过完全基于Web的界面在云上创建和部署自定义Galaxy实例。最后,我们将通过将开发的设施应用于现有的人类重新测序数据来测试我们的方法,以便在非常大的规模上发现导致人类遗传疾病的隐藏突变模式。 公共卫生相关性:越来越多的可用和廉价的高通量DNA测序为生物医学研究带来了巨大的希望,但信息学挑战阻碍了这种变革性技术潜力的充分实现。特别是生物医学研究人员的信息学和工程专业知识,以及足够的计算基础设施来分析这些巨大的数据集的可用性,限制了进展。该项目将解决这些问题,办法是将银河系统与“云计算”结合起来,前者是一个使复杂的计算分析可以利用和复制的系统,后者是一个按需购买计算资源的基础设施模式,使没有信息学专门知识的调查人员能够利用云资源进行数据密集型分析。
英文摘要
DESCRIPTION (provided by applicant): Project Summary Wide availability of "next-generation" sequencing (NGS) instruments has enabled any investigator, for a modest cost, to produce enormous amounts of DNA sequence data. However, working with these raw sequences presents significant problems for individual investigators, small labs, or core facilities. For an experimental group with no computational expertise, simply running a data analysis program is a barrier, let alone building a compute and data storage infrastructure capable of dealing with NGS data. Fortunately, a computational model - "Cloud computing" - has recently emerged and is ideally suited to the analysis of large- scale sequence data. In this model, computation and storage exist as virtual resources, which can be dynamically allocated and released as needed. Importantly, cloud resources can provide storage and computation at far less cost than dedicated resources for certain use cases. However, formidable challenges need to be addressed to make these resources available to individual investigators. Specifically, although cloud computing provides a way to acquire computational resources on demand, the resources provided are either virtual machines on the Internet or specific programming libraries, which are unusable for experimentalists. Thus, a viable analysis solution needs to be accessible and deployable without informatics expertise; it must efficiently and automatically use dynamically scalable resources, while taking into account time and cost; it must include appropriate analysis tools and easily support addition of new tools as they emerge. We have previously developed a software system - Galaxy (http://galaxyproject.org) - that provides a robust framework for addressing these needs. Here we propose to significantly extend this framework to allow any experimentalist to perform large-scale NGS analyses utilizing the power of cloud computing infrastructure. In particular, we will modify the existing Galaxy framework to run entirely within the cloud. We will adapt the way Galaxy schedules and executes jobs to make effective use of cloud-style. We will provide a mechanism for individual users to create and deploy custom Galaxy instances on a cloud through an entirely web-based interface. Finally, we will test our approach by applying the developed facilities to the existing human re- sequencing data in order to uncover hidden patters of mutations causing human genetic disease on a very large scale. PUBLIC HEALTH RELEVANCE: Project Narrative Increasingly available and inexpensive high-throughput DNA sequencing holds great promise for biomedical research, but informatics challenge block the full realization of the potential of this transformative technology. In particular progress is limited by the informatics and engineering expertise of biomedical researchers, and the availability of sufficient computational infrastructure to analyze these enormous datasets. This project will address these problems by bringing together Galaxy, a system for making complex computational analysis accessible and reproducible, with "cloud computing", an infrastructure model where computing resources are purchased on demand as needed, making it possible for investigators with no informatics expertise to perform data-intensive analysis using cloud resources.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Turning big data analysis infrastructure for HIV research
Tuning big data analysis infrastructure for HIV research
Tuning big data analysis infrastructure for HIV research
Democratization of Data Analysis in Life Sciences Through Galaxy
海外基金