Compressive Genomics for Large Omics Data Sets: Algorithms, Applications and Tools
Compressive Genomics for Large Omics Data Sets: Algorithms, Applications and Tools
批准号:
9546755
负责人:
BONNIE BERGER
金额:
$35.02万
依托单位国家:
美国
项目类别:
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-09-05 至 2020-08-31
关键词:
AccelerationAddressAdoptionAgeAlgorithmsAutistic DisorderAutoimmune DiseasesBioinformaticsBiologicalBiological ProcessBiologyBiomedical ResearchClinicalCloud ComputingCohort StudiesCollaborationsCommunitiesComplexComputer softwareComputersComputing MethodologiesDNA sequencingDataData CompressionData FilesData SetDevelopmentDimensionsDiseaseEnsureExhibitsFoundationsFractalsGenetic VariationGenomeGenomicsGoalsGrantHealthHumanIndianaIndividualIndustryInformaticsIntuitionInvestigationMainstreamingMalignant NeoplasmsMapsMetagenomicsMethodologyMolecularNew EnglandPatientsPatternPharmacogenomicsPrivacyProcessProgress ReportsResearchResearch PersonnelSavingsSecureSecuritySequence AnalysisStructureTechniquesTechnologyThe Cancer Genome AtlasTherapeuticTimeTranscriptUniversitiesVariantWorkautism spectrum disorderbasecloud platformcohortcomputer frameworkcomputerized toolscryptographydata exchangedata structuredesignencryptionfile formatgenomic datahuman DNAhuman RNA sequencinghuman datahuman diseaseimprovedinnovationinsightinterestmicrobiomemicrobiome therapeuticsmonomethoxypolyethylene glycolnext generation sequencingnovelsoftware developmenttooltransmission processwhole genome
中文摘要
项目摘要
高通量实验技术正在产生越来越大规模和复杂的基因组
序列数据集。虽然这些数据有希望揭示全新的生物学,但它们的纯粹
不确定性可能使它们的解释在计算上不可行。这个持续的目标
一个项目是设计和开发创新的基于压缩的算法技术,
处理大量的生物数据我们将在压缩搜索之外进行分支,
迫切需要在云中安全地存储和处理大规模基因组数据,
大规模宏基因组数据的洞察力。
关键的基本观察是,基因组数据是高度结构化的,表现出高度的
自相似性在我们之前的授予期间,我们利用了它的高冗余度和低分形
维度,以实现可扩展的压缩存储和加速序列数据的搜索,
以及与结构生物信息学和化学基因组学相关的其它生物数据类型。在这
更新,我们将继续利用结构(即,可压缩性),以:(i)
克服在共享敏感的人类数据(例如在云上)时出现的隐私问题;(ii)解决
新的挑战,超越搜索,与宏基因组数据;和(iii)寻求扩大采用
以前和新提出的压缩算法,用于工业,研究和临床使用。我们将
展示了我们的压缩技术对人类基因组和
宏基因组变异
我们将与合作伙伴Sahinalp的实验室(印第安纳州大学,布卢明顿)合作,
将这些工具应用于高通量数据集,包括自闭症谱系障碍(与艾萨克
Kohane和Evan Eichler)和癌症(与PCAWG,全基因组的泛癌症分析),
微生物组(与Eric Alm和Jian Peng)以及人类变异分析(GATK,与Eric
Lander和Eric Banks)。广泛的长期目标是将我们的压缩方法应用于
大量生物数据集可阐明仍然模糊的疾病分子格局。
这些目标的成功完成将导致计算方法和工具,
提高我们安全存储、访问和分析海量数据集的能力,
遗传变异的基本方面,以及实验的可检验的假设
调查事务所不仅所有开发的软件将公开提供,但作为我们的一部分,
为了实现整合目标,我们还将确保研究界可以利用我们的创新,
最小的努力通过我们的研究合作,我们将建立这些工具,并证明
它们与人类健康和疾病特征的相关性。
英文摘要
Project Summary
High-throughput experimental technologies are generating increasingly massive and complex genomic
sequence data sets. While these data hold the promise of uncovering entirely new biology, their sheer
enormity threatens to make their interpretation computationally infeasible. The continued goal of this
project is to design and develop innovative compression-based algorithmic techniques for efficiently
processing massive biological data. We will branch out beyond compressive search to address the
imminent need to securely store and process large-scale genomic data in the cloud, as well as to gain
insights from massive metagenomic data.
The key underlying observation is that genomic data is highly structured, exhibiting high degrees of
self-similarity. In our previous granting period, we exploited its high redundancy and low fractal
dimension to enable scalable compressive storage and acceleration for search of sequence data as
well as other biological data types relevant to structural bioinformatics and chemogenomics. In this
renewal, we will continue to capitalize on the structure (i.e., compressibility) of genomic data to: (i)
overcome privacy concerns that arise in sharing sensitive human data (e.g. on the cloud); (ii) address
new challenges, beyond search, with metagenomic data; and (iii) seek to widen the adoption of the
previous and newly-proposed compressive algorithms for industry, research, and clinical use. We will
demonstrate the utility of our compressive techniques to the characterization of human genomic and
metagenomic variation.
We will collaborate with co-I Sahinalp's lab (Indiana University, Bloomington) on developing and
applying these tools to high-throughput data sets including autism spectrum disorder (with Isaac
Kohane and Evan Eichler) and cancer (with PCAWG, Pan Cancer Analysis of Whole Genomes), the
microbiome (with Eric Alm and Jian Peng), as well as human variation analysis (GATK, with Eric
Lander and Eric Banks). The broad, long-term goal is to apply our compressive approach to
massive biological data sets to elucidate the still obscure molecular landscape of diseases.
Successful completion of these aims will result in computational methods and tools that will significantly
increase our ability to securely store, access and analyze massive data sets and will reveal
fundamental aspects of genetic variation, as well as testable hypotheses for experimental
investigations. Not only will all developed software be made publicly available, but as part of our
integration aim, we will also ensure that the research community can make use of our innovations with
minimal effort. Through our research collaborations, we will both build these tools and demonstrate
their relevance to the characterization of human health and disease.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Manifold representations and active learning for 21 st century biology
-
批准号:10401890
-
项目类别:
-
资助金额:$35.99万
-
财政年份:2021
-
负责人:BONNIE BERGER
-
依托单位:
Manifold representations and active learning for 21 st century biology
-
批准号:10207091
-
项目类别:
-
资助金额:$47.83万
-
财政年份:2021
-
负责人:BONNIE BERGER
-
依托单位:
Manifold representations and active learning for 21 st century biology
-
批准号:10670057
-
项目类别:
-
资助金额:$35.89万
-
财政年份:2021
-
负责人:BONNIE BERGER
-
依托单位:
Developing high-throughput genetic perturbation strategies for single cells in cancer organoids
-
批准号:10004966
-
项目类别:
-
资助金额:$92.22万
-
财政年份:2020
-
负责人:BONNIE BERGER
-
依托单位:
Privacy-preserving genomic medicine at scale
-
批准号:10266081
-
项目类别:
-
资助金额:$67.49万
-
财政年份:2020
-
负责人:BONNIE BERGER
-
依托单位:
Privacy-preserving genomic medicine at scale
-
批准号:10459604
-
项目类别:
-
资助金额:$66.28万
-
财政年份:2020
-
负责人:BONNIE BERGER
-
依托单位:
Privacy-preserving genomic medicine at scale
-
批准号:10662349
-
项目类别:
-
资助金额:$66.75万
-
财政年份:2020
-
负责人:BONNIE BERGER
-
依托单位:
Developing high-throughput genetic perturbation strategies for single cells in cancer organoids
-
批准号:10212991
-
项目类别:
-
资助金额:$92.22万
-
财政年份:2020
-
负责人:BONNIE BERGER
-
依托单位:
Compressive genomics for large omics data sets: Algorithms applications & tools
-
批准号:8849927
-
项目类别:
-
资助金额:$20.94万
-
财政年份:2013
-
负责人:BONNIE BERGER
-
依托单位:
Compressive genomics for large omics data sets: Algorithms applications & tools
-
批准号:8599836
-
项目类别:
-
资助金额:$21.79万
-
财政年份:2013
-
负责人:BONNIE BERGER
-
依托单位:
Compressive Genomics for Large Omics Data Sets: Algorithms, Applications and Tools
-
批准号:9247325
-
项目类别:
-
资助金额:$37.2万
-
财政年份:2013
-
负责人:BONNIE BERGER
-
依托单位:
Compressive genomics for large omics data sets: Algorithms applications & tools
-
批准号:8730209
-
项目类别:
-
资助金额:$21.32万
-
财政年份:2013
-
负责人:BONNIE BERGER
-
依托单位:
MIT/Whitehead/Broad Computational Genetics Training Program
-
批准号:8132612
-
项目类别:
-
资助金额:$17.29万
-
财政年份:2009
-
负责人:BONNIE BERGER
-
依托单位:
Structure-Based Prediction of the Interactome
-
批准号:7895360
-
项目类别:
-
资助金额:$28.99万
-
财政年份:2009
-
负责人:BONNIE BERGER
-
依托单位:
Structure-Based Prediction of the Interactome
-
批准号:8054929
-
项目类别:
-
资助金额:$28.99万
-
财政年份:2008
-
负责人:BONNIE BERGER
-
依托单位:
Structure based prediction of the interactome
-
批准号:8439763
-
项目类别:
-
资助金额:$32.03万
-
财政年份:2008
-
负责人:BONNIE BERGER
-
依托单位:
Structure based prediction of the interactome
-
批准号:8848384
-
项目类别:
-
资助金额:$32.08万
-
财政年份:2008
-
负责人:BONNIE BERGER
-
依托单位:
Structure-Based Prediction of the Interactome
-
批准号:7464375
-
项目类别:
-
资助金额:$28.09万
-
财政年份:2008
-
负责人:BONNIE BERGER
-
依托单位:
Structure based Prediction of the interactome
-
批准号:9549093
-
项目类别:
-
资助金额:$34.65万
-
财政年份:2008
-
负责人:BONNIE BERGER
-
依托单位:
Structure-Based Prediction of the Interactome
-
批准号:7797574
-
项目类别:
-
资助金额:$30.12万
-
财政年份:2008
-
负责人:BONNIE BERGER
-
依托单位:
海外基金