Relationship inference in large genetic data
Relationship inference in large genetic data
批准号:
9076754
负责人:
WEI-MIN CHEN
金额:
$39.5万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-07-21 至 2021-05-31
关键词:
AdmixtureAlgorithmsChromosome MappingCollectionComputer softwareCustomDataData AnalysesData SetDetectionDevelopmentDistantFamilyFamily RelationshipFundingFutureGeneticGenetic studyGenomeGenotypeGrantHigh-Throughput Nucleotide SequencingInbreedingIndividualInsulin-Dependent Diabetes MellitusMapsMethodsNational Heart, Lung, and Blood InstituteParentsPhenotypePlayPopulationProceduresQuality ControlResearch PersonnelResolutionRoleSample SizeSamplingSoftware ToolsStatistical MethodsStructureTechnologyTestingUnited States National Institutes of HealthVariantVertebral columnbasecomputerized toolsexomeexome sequencingfamily structuregenetic pedigreegenome sequencinggenomic datahuman diseaseoffspringpublic health relevancerare variantreconstructiontoolwhole genome
中文摘要
描述(由申请人提供):高通量测序和定制基因分型阵列的技术进步使基因研究比以往任何时候都更大。在过去的几年里,产生全基因组测序(WGS)数据的研究数量大幅增加,仅NHLBI一项就有望在明年产生超过30,000人的WGS数据。目前正在收集的外显子组序列数据现在接近200,000个对象,未来NIH和私人资助的项目将很快产生具有类似样本量的WGS。人们迫切需要充分利用新产生的大量数据,包括更好地识别和利用关联性信息。由于它的计算效率,我们的关系推理工具(King)在过去几年里一直是大型遗传研究中推断关系的主要软件工具。随着高通量基因分型、外显子组和全基因组序列数据带来的挑战和巨大的机遇,迫切需要一种更快、更可靠、更强大的关系推理程序和工具。这样一种工具将开启一种可能性,在目前可用的方法之外,通知罕见的变异关联。我们建议开发稳健和计算高效的算法来推断由1,000-100,000个人组成的大数据集中的亲缘关系和远亲关系。快速算法将允许识别
在由100,000个个体组成的大型数据集中的密切关系,以及专门为来自WGS技术的罕见变异数据提出的算法,将使我们能够在近亲繁殖、种群结构(包括种群混合)和/或样本污染存在的情况下,以及在更高的程度上更可靠地推断关系。此外,我们计划开发一个基于我们的快速关系推理算法的集成工具集,例如谱系重建、质量控制(QC)和基于家庭的关联方法。初步的数据分析表明,我们的算法可以在12秒内识别出1000个基因组数据中的所有亲密关系。我们还成功地推断出了一个只包含远亲(2级和3级)关系的扩展家系,代表了一位阿姨、她的侄女和她的第一个表亲。我们建议的方法将在自由分布软件(King)中实现,允许其他研究人员直接将这些方法应用于他们自己的测序和其他高通量阵列数据的分析。我们预计,在未来几年,这里开发的关系推理方法将在大规模遗传/基因组数据的质量控制和分析中发挥重要作用。
英文摘要
DESCRIPTION (provided by applicant): Technological advances in high-throughput sequencing and custom genotyping arrays are making genetic studies larger than ever. The number of studies generating whole genome sequencing (WGS) data has increased substantially over the past few years, and the NHLBI alone is expected to generate WGS data on over ~30,000 individuals in the next year. Ongoing collections of exome sequence data are now approaching 200,000 subjects, and future NIH- and private-funded projects will soon generate WGS with similar sample size. There is a great need to make full use of the large amount of newly generated data, including a better way to identify and utilize relatedness information. Due to its computational efficiency, our relationship inference tool (KING) has been the main software tool to infer relationships in large genetic studies in the past few years. With the challenges and great opportunities provided by high-throughput genotyping, exome and whole genome sequence data, an even faster, more reliable and more powerful relationship inference procedure and tool is urgently needed. Such a tool would open possibilities to inform rare variant association beyond currently available approaches. We propose to develop robust and computationally efficient algorithms to infer close and distant family relationships in large datasets consisting of 1,000s-100,000s of individuals. The fast algorithm will allow identification
of close relationships in large datasets consisting of >100,000 individuals, and the algorithms that are proposed specifically for the rare variant data from the WGS technology will allow us to infer relationships more reliably in the presence of inbreeding, population structure (including population admixture), and/or sample contamination, and also at a higher-order of degree. Further, we plan to develop an integrated toolset that is based on our fast relationship inference algorithms, such as pedigree reconstruction, Quality Control (QC), and family-based association methods. Preliminary data analysis shows our algorithm can identify all close relationships in the 1000 Genomes data in 12 seconds. We also successfully inferred an extended pedigree containing only distant (2nd- and 3rd-degree) relationships representing an aunt, her niece, and her first cousin. Our proposed methods will be implemented in freely distributed software (KING), allowing other investigators to apply the methods directly to analysis of their own sequencing and other high-throughput array data. We expect the relationship inference methods developed here will play an important role in the quality control and analysis of large sets of genetic/genomic data in the coming years.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Family-based rare variant association methods for quantitative traits
-
批准号:8514675
-
项目类别:
-
资助金额:$7.9万
-
财政年份:2012
-
负责人:WEI-MIN CHEN
-
依托单位:
Family-based rare variant association methods for quantitative traits
-
批准号:8355029
-
项目类别:
-
资助金额:$7.9万
-
财政年份:2012
-
负责人:WEI-MIN CHEN
-
依托单位:
海外基金