Integrated variation detection annotation and analysis for high-throughout seque
Integrated variation detection annotation and analysis for high-throughout seque
批准号:
8220672
负责人:
Kai Wang
金额:
$36.01万
依托单位国家:
美国
项目类别:
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-03-23 至 2017-02-28
关键词:
AlgorithmsBase SequenceBenchmarkingBiologicalCategoriesCodeCommunitiesComputational algorithmComputer softwareComputing MethodologiesCopy Number PolymorphismDNA SequenceDataData AnalysesData SourcesDatabasesDetectionDevelopmentDiseaseDisease susceptibilityFunctional RNAGene FrequencyGenesGenetic VariationGenomeGenomicsGenotypeGleanHumanInformaticsKnowledgeLeadLocationMetabolic PathwayMethodsModelingOntologyPathway AnalysisPathway interactionsPhenotypePopulationPropertyRNAReadingRegulator GenesResearch PersonnelResourcesRoleSamplingScientific Advances and AccomplishmentsScoring MethodSequence AlignmentSoftware ToolsSourceStatistical MethodsSystemTestingUpdateVariantWorkbasedata formatdosageexperiencegenetic variantgenome-widehigh throughput analysishuman subjectimprovedmarkov modelmethod developmentopen sourcepopulation basedsimulationtechnique developmenttoolvector
中文摘要
描述(申请人提供):关于各种物种基因组的高通量测序(HTS)数据正在以前所未有的速度产生。然而,处理这些数据的计算和统计方法的发展滞后,造成了正在产生的海量数据与可以收集的生物学知识之间的差距。在这里,我们建议开发一个用于HTS数据的遗传变异检测、注释和分析的集成系统,从而缩小社区面临的关键差距。在目标1中,我们将开发一种基于隐马尔可夫模型(HMM)的计算算法,该算法结合了多个信息源,包括序列深度、等位基因剂量、群体等位基因频率和成对末端阅读距离,以可靠而有效地检测拷贝数变异(CNV)。考虑到大量的SNPs、Indels和CNV,研究人员面临着识别功能上重要的变体的子集的挑战。在目标2中,我们将开发一个全面的功能注释管道,利用来自许多大型基因组学项目的数据库信息,注释编码和非编码变体的功能重要性,并为每个变体生成“功能向量”。这些功能载体可以帮助生物学家解释测序结果,并帮助统计遗传学家利用测序数据开发知情的关联测试。需要适当的统计方法来分析群体水平的测序数据,以便识别可能导致疾病易感性或表型变异的基因组变异。在目标3中,我们将开发一种分层建模策略,利用每个变体的功能载体信息,对基因、基因组区域或生物路径进行关联测试,如本体类别和基因调节/代谢路径。最后,在目标4中,我们将通过模拟和真实数据分析来测试每种方法的性能,并开发、分发和支持实现所提出方法的免费可用的软件包。我们相信,记录良好和得到支持的软件实施将使其他研究人员能够从该项目所产生的方法和科学进步中获得最大的信息。AIMS的成功完成将使研究人员能够充分调查已经或将产生的大量测序数据,从而有助于我们了解遗传变异如何影响表型可变性。
公共卫生相关性:尽管高通量测序(HTS)技术迅速进步,但处理这些数据的计算和统计方法的开发滞后,造成了正在产生的海量数据与可以收集的生物学知识之间的差距。在这里,我们建议开发一个集成的系统来检测变异体,注释变异体并分析它们的基因型-表型关联。AIMS的成功完成将使研究人员能够充分调查已经或将产生的大量测序数据,从而有助于我们了解遗传变异如何影响表型可变性。
英文摘要
DESCRIPTION (provided by applicant): High-throughput sequencing (HTS) data on the genomes of a diverse number of species are being produced at an unprecedented rate. However, the development of computational and statistical approaches for handling these data lags behind, creating a gap between the massive data being generated and the biological knowledge that could be gleaned. Here we propose to develop an integrated system for genetic variation detection, annotation and analysis for HTS data, therefore reducing the critical gap faced by the community. In Aim 1, we will develop a hidden Markov model (HMM) based computational algorithm that incorporates multiple sources of information, including sequence depth, allelic dosage, population allele frequency and paired-end reads distance, for reliable yet efficient detection of copy number variations (CNVs). Given a large list of SNPs, indels and CNVs, researchers are faced with the challenge of identifying a subset of functionally important variants. In Aim 2, we will develop a comprehensive functional annotation pipeline to annotate functional importance of coding and non-coding variants, utilizing database information from many large-scale genomics projects, and generate a "functional vector" for each variant. These functional vectors can help biologists interpret sequencing results and help statistical geneticists develop informed association tests using sequencing data. Appropriate statistical methods are needed to analyze population-level sequencing data, in order to identify genomic variants that may contribute to disease susceptibility or phenotypic variability. In Aim 3, we will develop a hierarchical modeling strategy, which utilizes functional vector information for each variant, to perform association tests on genes, genomic regions, or biological pathways, such as ontology categories and gene regulatory/metabolic pathways. Finally, in Aim 4, we will test the properties of each approach via simulation and real data analysis, and develop, distribute and support freely available software packages implementing the proposed methods. We believe that well-documented and supported software implementations will allow other researchers to yield the maximum information from the methodological and scientific advances that result from this project. Successful completion of the aims will enable researchers to fully investigate the massive amounts of sequencing data that have been or will be generated, thus contributing to our understanding on how genetic variants influence phenotype variability.
PUBLIC HEALTH RELEVANCE: Despite the rapid advancement of high-throughput sequencing (HTS) techniques, the development of computational and statistical approaches for handling these data lags behind, creating a gap between the massive data being generated and the biological knowledge that could be gleaned. Here we propose to develop an integrated system to detect variants, annotate variants and analyze them for genotype-phenotype associations. Successful completion of the aims will enable researchers to fully investigate the massive amounts of sequencing data that have been or will be generated, thus contributing to our understanding on how genetic variants influence phenotype variability.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Dietary prevention for colorectal cancer: targeting the bile acid/gut microbiome axis
-
批准号:10723195
-
项目类别:
-
资助金额:$12.41万
-
财政年份:2023
-
负责人:Kai Wang
-
依托单位:
Novel bioinformatics methods to detect DNA and RNA modifications using Nanopore long-read sequencing
-
批准号:10792416
-
项目类别:
-
资助金额:$70.96万
-
财政年份:2023
-
负责人:Kai Wang
-
依托单位:
Improving chemical exposome target prediction by application of Coupled Matrix/Tensor-Matrix/Tensor Completion algorithms
-
批准号:10734136
-
项目类别:
-
资助金额:$11.8万
-
财政年份:2023
-
负责人:Kai Wang
-
依托单位:
Detection and annotation of structural variants from long-read sequencing
-
批准号:10378720
-
项目类别:
-
资助金额:$44.0万
-
财政年份:2019
-
负责人:Kai Wang
-
依托单位:
Integrated Variation Detection Annotation and Analysis
-
批准号:9402354
-
项目类别:
-
资助金额:$22.73万
-
财政年份:2016
-
负责人:Kai Wang
-
依托单位:
UNDERSTANDING THE FUNCTIONAL IMPACTS OF GENETIC VARIANTS IN MENTAL DISORDERS
-
批准号:9389287
-
项目类别:
-
资助金额:$38.15万
-
财政年份:2016
-
负责人:Kai Wang
-
依托单位:
Role of MTA3 in trophoblast function and placental development
-
批准号:8919934
-
项目类别:
-
资助金额:$7.48万
-
财政年份:2014
-
负责人:Kai Wang
-
依托单位:
Integrated variation detection annotation and analysis for high-throughout seque
-
批准号:8448070
-
项目类别:
-
资助金额:$34.46万
-
财政年份:2012
-
负责人:Kai Wang
-
依托单位:
Integrated variation detection annotation and analysis for high-throughout seque
-
批准号:8813611
-
项目类别:
-
资助金额:$35.36万
-
财政年份:2012
-
负责人:Kai Wang
-
依托单位:
Integrated variation detection annotation and analysis for high-throughout seque
-
批准号:8628856
-
项目类别:
-
资助金额:$35.43万
-
财政年份:2012
-
负责人:Kai Wang
-
依托单位:
BUILDING A HIGH PERFORMANC BIOINFORMATICS CLUSTER
-
批准号:7610346
-
项目类别:
-
资助金额:$2.91万
-
财政年份:2007
-
负责人:Kai Wang
-
依托单位:
海外基金