High-performance mixed model toolset for integrative omics analysis of big data
High-performance mixed model toolset for integrative omics analysis of big data
批准号:
9312511
负责人:
JEFFREY R O'CONNELL
金额:
$58.48万
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-04-15 至 2020-03-31
关键词:
AlgorithmsAttentionBedsBig DataBiologicalBiological ModelsBiologyCloud ComputingCodeCollaborationsCommunitiesComplexComplex Genetic TraitComputer softwareDataData AnalysesData SetDevelopmentDiseaseEpigenetic ProcessEquationFamilyGeneticGenetic EpistasisGenomicsGenotypeGoalsHealthHeterogeneityHumanInvestmentsMemoryMeta-AnalysisMethodologyMitochondriaModelingNational Heart, Lung, and Blood InstituteNational Human Genome Research InstitutePerformancePhasePhenotypePlayPopulationProcessPublishingResearchResearch PersonnelResource AllocationResourcesRoleSample SizeSamplingSequence AnalysisShapesSystemSystems BiologyTechnologyTestingTimeTrans-Omics for Precision MedicineVariantWeightWorkanalytical methodanimal breedingbasebiological systemscloud basedcohortcostdata accessdata managementdata spaceepigenome-wide association studiesfile formatflexibilitygenetic analysisgenetic pedigreegenomic dataimprovedinsightlarge scale productionmethod developmentnovelnovel strategiesprecision medicinepressurerare variantresponsescale upsimulationsimulation softwareterabytetooltraitvirtualwhole genomeworking group
中文摘要
点击翻译按钮获取中文摘要
英文摘要
PROJECT SUMMARY/ABSTRACT
The recent large scale production of whole genome sequence and other multi-omics in TOPMed and other
projects calls for parallel development of comprehensive, powerful and flexible toolset capable of large data
management, analysis and integration. Mega/integrated analyses are essential to fully utilize these data to
elucidate the complexity of the biological mechanisms and advance our understanding of complex trait biology to
drive precision medicine. TOPMed estimates that the VCF for 60,000 subjects will contain 400M variants and
require 100TB of space, and much of our current genetic analysis toolset does not scale up to these data sizes.
For rare variant analysis, mixed model mega analysis is more powerful than meta-analysis as mega analysis can
include additional random effects to account for genetic relatedness between all subjects and cross-study
phenotypic, genetic and environmental heterogeneity. However cross-study mega analysis within the mixed
model is still an uncharted territory. We believe mega analysis will spur more creative analysis approaches
provided the needed toolsets are available. In cloud computing “time is money”, and new approaches are
required to solve structural differences in resource allocation and data access compared to local computing.
MMAP (Mixed Models for Analysis of Pedigrees/Populations) is robust mixed model software that already
published mixed model analysis on a sample size of 90,000 that included dominance variance and developed a
cloud-efficient version of mixed model rare variant analysis. The goal of this proposal is to further expand and
improve this toolset to deliver to the research community a flexible, versatile, and comprehensive cross-platform
mixed model toolset scalable to efficient local and cloud analysis of large WGS and omics data. We plan to
implement several new features in our toolset including: 1) Efficient binary genotype file format for optimal
storage of terabyte VCF genotypes. 2) Large-scale modeling of non-additive variation such as dominance, X-
lined, mitochondrial and epistasis. 3) Optimized rare variant analysis with flexible integration of annotation and
variant weighting resources. 4) Optimized expression/epigenome-wide association (EWA) analysis. 5)
Comprehensive multi-omics integration into the mixed model as fixed and random effects. 6) Development of a
multi-omics simulation software to guide systems biology modeling. 7) Integrating mixed model equations for
prediction from animal breeding. This proposal will deliver the research community an analysis toolset that will
push research boundaries well beyond additive SNP association to a space filled with complex biological fixed
and random effects models integrating the full spectrum of multi-omics data. We plan to develop a multi-omics
simulation tool to better understand the complex evolutionary processes that shape the complex trait landscape.
Our toolset will be extensively shaped by collaboration with TOPMed working groups to meet analysis priorities
and develop analysis plans. Our toolset will surely evolve in novel and unexpected directions in response to new
ideas and challenges as we dive deeper into this unique data set.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Elucidating the ancestry-specific genetic and environmental architecture of cardiometabolic traits across All of Us ethnic groups
-
批准号:10796028
-
项目类别:
-
资助金额:$19.31万
-
财政年份:2023
-
负责人:JEFFREY R O'CONNELL
-
依托单位:
Genome-wide Association in Families: Data Integrity, Design and Methods Issue
-
批准号:7104529
-
项目类别:
-
资助金额:$30.59万
-
财政年份:2006
-
负责人:JEFFREY R O'CONNELL
-
依托单位:
Genome-wide Association in Families: Data Integrity, Design and Methods Issue
-
批准号:7246523
-
项目类别:
-
资助金额:$28.98万
-
财政年份:2006
-
负责人:JEFFREY R O'CONNELL
-
依托单位:
Genome-wide Association in Families: Data Integrity, Design and Methods Issue
-
批准号:7421072
-
项目类别:
-
资助金额:$28.41万
-
财政年份:2006
-
负责人:JEFFREY R O'CONNELL
-
依托单位:
RAPID MULTIPOINT METHODS FOR MAPPING COMPLEX DISEASES
-
批准号:2864800
-
项目类别:
-
资助金额:$16.73万
-
财政年份:1998
-
负责人:JEFFREY R O'CONNELL
-
依托单位:
RAPID MULTIPOINT METHODS FOR MAPPING COMPLEX DISEASES
-
批准号:6043142
-
项目类别:
-
资助金额:$20.37万
-
财政年份:1998
-
负责人:JEFFREY R O'CONNELL
-
依托单位:
RAPID MULTIPOINT METHODS FOR MAPPING COMPLEX DISEASES
-
批准号:6169588
-
项目类别:
-
资助金额:$20.98万
-
财政年份:1998
-
负责人:JEFFREY R O'CONNELL
-
依托单位:
国内基金
海外基金
多模态超声VisTran-Attention网络评估早期子宫颈癌保留生育功能手术可行性
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:郑巧
-
依托单位:
Ultrasomics-Attention孪生网络早期精准评估肝内胆管癌免疫治疗的研究
-
批准号:--
-
项目类别:面上项目
-
资助金额:52万元
-
批准年份:2022
-
负责人:陈立达
-
依托单位: