课题基金 / 基金详情

Compressive Genomics for Large Omics Data Sets: Algorithms, Applications and Tools

Compressive Genomics for Large Omics Data Sets: Algorithms, Applications and Tools
大型组学数据集的压缩基因组学:算法、应用程序和工具
批准号:
9546755
负责人:
BONNIE BERGER
金额:
$35.02万
依托单位国家:
美国
项目类别:
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-09-05 至 2020-08-31

项目摘要

项目成果

BONNIE BERGER的其他基金

相似基金

相关文献

中文摘要
翻译
项目总结: -- 高通量实验技术正在产生越来越庞大和复杂的基因组 序列数据集。虽然这些数据有望发现全新的生物学,但它们的纯粹 巨大的威胁性使得他们的解释在计算上是不可行的。这样做的持续目标是 该项目是设计和开发创新的基于压缩的算法技术,以有效地 处理海量的生物数据。我们将扩展到压缩搜索之外,以解决 迫切需要在云中安全地存储和处理大规模基因组数据,以及获得 来自海量元基因组数据的见解。-- -- 关键的潜在观察是基因组数据是高度结构化的,表现出高度的 自相似。在我们之前的授予期间,我们利用了它的高冗余性和低分形性 支持可扩展压缩存储和加速的维度,用于搜索序列数据 以及与结构生物信息学和化学基因组学相关的其他生物数据类型。在这 更新,我们将继续利用基因组数据的结构(即可压缩性)来:(I) 克服在共享敏感人类数据(例如在云上)时出现的隐私问题;(Ii)解决 在搜索之外,对元基因组数据提出新的挑战;以及(3)寻求扩大采用 以前和最近提出的用于工业、研究和临床使用的压缩算法。我们会 展示我们的压缩技术在表征人类基因组和 元基因组变异。 -- 我们将与co-I Sahinalp的实验室(印第安纳大学,布鲁明顿)合作开发和 将这些工具应用于高通量数据集,包括自闭症谱系障碍(与Isaac Kohane和Evan Eichler)和癌症(PCAWG,全基因组的泛癌分析), 微生物组(与Eric Alm和Jian Peng),以及人类变异分析(GATK,与Eric 兰德和埃里克·班克斯)。广泛而长期的目标是将我们的综合方法应用于 大量的生物学数据集来阐明仍然鲜为人知的疾病分子图景。 -- 这些目标的成功实现将导致计算方法和工具的显著提高 提高我们安全存储、访问和分析海量数据集的能力,并将揭示 遗传变异的基本方面,以及实验的可检验假说 调查。所有开发的软件不仅将公开提供,而且将作为我们的 整合的目标,我们还将确保研究界能够利用我们的创新与 最小的努力。通过我们的研究合作,我们将构建这些工具并演示 它们与人类健康和疾病的特征有关。
英文摘要
Project Summary    High-throughput experimental technologies are generating increasingly massive and complex genomic sequence data sets. While these data hold the promise of uncovering entirely new biology, their sheer enormity threatens to make their interpretation computationally infeasible. The continued goal of this project is to design and develop innovative compression-based algorithmic techniques for efficiently processing massive biological data. We will branch out beyond compressive search to address the imminent need to securely store and process large-scale genomic data in the cloud, as well as to gain insights from massive metagenomic data.     The key underlying observation is that genomic data is highly structured, exhibiting high degrees of self-similarity. In our previous granting period, we exploited its high redundancy and low fractal dimension to enable scalable compressive storage and acceleration for search of sequence data as well as other biological data types relevant to structural bioinformatics and chemogenomics. In this renewal, we will continue to capitalize on the structure (i.e., compressibility) of genomic data to: (i) overcome privacy concerns that arise in sharing sensitive human data (e.g. on the cloud); (ii) address new challenges, beyond search, with metagenomic data; and (iii) seek to widen the adoption of the previous and newly-proposed compressive algorithms for industry, research, and clinical use. We will demonstrate the utility of our compressive techniques to the characterization of human genomic and metagenomic variation.   We will collaborate with co-I Sahinalp's lab (Indiana University, Bloomington) on developing and applying these tools to high-throughput data sets including autism spectrum disorder (with Isaac Kohane and Evan Eichler) and cancer (with PCAWG, Pan Cancer Analysis of Whole Genomes), the microbiome (with Eric Alm and Jian Peng), as well as human variation analysis (GATK, with Eric Lander and Eric Banks). The broad, long-term goal is to apply our compressive approach to massive biological data sets to elucidate the still obscure molecular landscape of diseases.    Successful completion of these aims will result in computational methods and tools that will significantly increase our ability to securely store, access and analyze massive data sets and will reveal fundamental aspects of genetic variation, as well as testable hypotheses for experimental investigations. Not only will all developed software be made publicly available, but as part of our integration aim, we will also ensure that the research community can make use of our innovations with minimal effort. Through our research collaborations, we will both build these tools and demonstrate their relevance to the characterization of human health and disease.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Manifold representations and active learning for 21 st century biology
Manifold representations and active learning for 21 st century biology
Manifold representations and active learning for 21 st century biology
Developing high-throughput genetic perturbation strategies for single cells in cancer organoids
海外基金