Powerful and Adaptive Statistical Methods for Sequencing Studies
Powerful and Adaptive Statistical Methods for Sequencing Studies
批准号:
9344676
负责人:
Xinge Jessie Jeng
金额:
$7.58万
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-02 至 2018-08-31
关键词:
AddressArchivesBiologicalComplexComputing MethodologiesCopy Number PolymorphismDataData AnalysesData SetDetectionDevelopmentDimensionsDiseaseFamilial diseaseFollow-Up StudiesGene FrequencyGenesGeneticGenetic VariationGenomic SegmentGenomicsGoalsHeritabilityIndividualInheritedInvestigationKnowledgeLanguageLarge-Scale SequencingMental disordersMethodsMinorModelingMolecularMutationNatural SelectionsNucleotidesPathway interactionsProbabilityProceduresProgramming LanguagesPropertyResearch PersonnelResolutionResortRoleSample SizeSamplingScienceScreening procedureSignal TransductionSourceStatistical MethodsTechniquesTechnologyTestingThe Cancer Genome AtlasUrsidae FamilyVariantbasecomputerized toolsfollow-upgenetic variantgenome wide association studygenomic variationhigh dimensionalitynext generation sequencingnovelprogramspublic health relevancerare variantrepositoryscreeningsimulationstatisticssuccesstheoriestooltrait
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): Next-generation sequencing (NGS) data are being increasingly generated over the last few years. Encompassing the full spectrum of genomic variations, they hold the promise of identifying new sources of heritability from rare variants tha were eluded in traditional genome-wide association (GWA) studies. Despite substantial progresses in recent years, current methods are, nonetheless, limited in terms of power and robustness towards the analysis of NGS data that are characterized by extreme high dimensionality and low minor allele frequency (MAF). New methods are needed to adapt to these statistical challenges in order to achieve the full potential of NGS data in identifying genetic variations contributing to missing disease heritability. The goal of this project is to develop powerful and adaptive statistical methods for the analysis of sequencing studies. Specifically, the project aims to (1) develop an adaptive variants screening procedure that can efficiently account for a large proportion of causal rare variants while significantly reducing te data dimension for follow-up analysis; and to (2) provide an objective procedure for samples-size calculation to direct follow-up studies and to pinpoint the causal variants with high confidence. The proposed procedures are very general and can accommodate a wide spectrum of models, test statistics, and data scenarios. They are completely data-driven and can automatically adapt to the underlying sparsity of the data. Moreover, the proposed methods are computationally efficient under extreme high dimensionality. These desirable properties make the proposed methods applicable to a myriad of high-dimensional applications. Rigorous theory will be developed to understand the role of sparsity and extreme high dimensionality in NGS data analysis, and comprehensive simulations will be performed to study the proposed methods. In addition, this project will provide computationally efficient programs and evaluate the methods using several recent NGS datasets. The programs will be developed in R and efficient Fortran languages. Our computational package will be made publicly available to allow investigators to apply our procedures widely in sequencing studies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金