Relationship inference in large genetic data
Relationship inference in large genetic data
批准号:
9076754
负责人:
WEI-MIN CHEN
金额:
$39.5万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-07-21 至 2021-05-31
关键词:
AdmixtureAlgorithmsChromosome MappingCollectionComputer softwareCustomDataData AnalysesData SetDetectionDevelopmentDistantFamilyFamily RelationshipFundingFutureGeneticGenetic studyGenomeGenotypeGrantHigh-Throughput Nucleotide SequencingInbreedingIndividualInsulin-Dependent Diabetes MellitusMapsMethodsNational Heart, Lung, and Blood InstituteParentsPhenotypePlayPopulationProceduresQuality ControlResearch PersonnelResolutionRoleSample SizeSamplingSoftware ToolsStatistical MethodsStructureTechnologyTestingUnited States National Institutes of HealthVariantVertebral columnbasecomputerized toolsexomeexome sequencingfamily structuregenetic pedigreegenome sequencinggenomic datahuman diseaseoffspringpublic health relevancerare variantreconstructiontoolwhole genome
中文摘要
描述(由申请人提供):高通量测序和定制基因分型阵列的技术进步使遗传学研究比以往任何时候都更大。在过去的几年里,产生全基因组测序(WGS)数据的研究数量大幅增加,预计明年仅NHLBI就将产生超过30,000人的WGS数据。目前正在收集的外显子组序列数据接近20万个受试者,未来NIH和私人资助的项目将很快产生具有类似样本量的WGS。非常需要充分利用大量新产生的数据,包括更好地确定和利用相关性信息。由于其计算效率,我们的关系推理工具(KING)在过去几年中一直是大型遗传研究中推断关系的主要软件工具。随着高通量基因分型、外显子组和全基因组序列数据提供的挑战和巨大机遇,迫切需要更快、更可靠和更强大的关系推理程序和工具。这样的工具将开辟可能性,以告知罕见的变异关联超出目前可用的方法。 我们建议开发强大的和计算效率高的算法来推断由1,000 - 100,000个个体组成的大型数据集中的亲密和疏远的家庭关系。快速算法将允许识别
在由> 100,000个个体组成的大型数据集中,紧密关系的最大化,并且专门针对来自WGS技术的罕见变异数据提出的算法将允许我们在存在近亲繁殖、群体结构(包括群体混合)和/或样品污染的情况下更可靠地推断关系,并且还在更高的程度上。此外,我们计划开发一个集成的工具集,该工具集基于我们的快速关系推理算法,如谱系重建,质量控制(QC)和基于家庭的关联方法。初步数据分析表明,我们的算法可以在12秒内识别1000个基因组数据中的所有密切关系。我们还成功地推断出一个扩展的谱系,其中只包含代表阿姨、侄女和堂兄弟姐妹的远亲(第二和第三度)关系。我们提出的方法将在自由分发的软件(KING)中实现,允许其他研究人员直接应用这些方法来分析他们自己的测序和其他高通量阵列数据。我们希望在这里开发的关系推理方法将在质量控制和分析的大型遗传/基因组数据在未来几年中发挥重要作用。
英文摘要
DESCRIPTION (provided by applicant): Technological advances in high-throughput sequencing and custom genotyping arrays are making genetic studies larger than ever. The number of studies generating whole genome sequencing (WGS) data has increased substantially over the past few years, and the NHLBI alone is expected to generate WGS data on over ~30,000 individuals in the next year. Ongoing collections of exome sequence data are now approaching 200,000 subjects, and future NIH- and private-funded projects will soon generate WGS with similar sample size. There is a great need to make full use of the large amount of newly generated data, including a better way to identify and utilize relatedness information. Due to its computational efficiency, our relationship inference tool (KING) has been the main software tool to infer relationships in large genetic studies in the past few years. With the challenges and great opportunities provided by high-throughput genotyping, exome and whole genome sequence data, an even faster, more reliable and more powerful relationship inference procedure and tool is urgently needed. Such a tool would open possibilities to inform rare variant association beyond currently available approaches. We propose to develop robust and computationally efficient algorithms to infer close and distant family relationships in large datasets consisting of 1,000s-100,000s of individuals. The fast algorithm will allow identification
of close relationships in large datasets consisting of >100,000 individuals, and the algorithms that are proposed specifically for the rare variant data from the WGS technology will allow us to infer relationships more reliably in the presence of inbreeding, population structure (including population admixture), and/or sample contamination, and also at a higher-order of degree. Further, we plan to develop an integrated toolset that is based on our fast relationship inference algorithms, such as pedigree reconstruction, Quality Control (QC), and family-based association methods. Preliminary data analysis shows our algorithm can identify all close relationships in the 1000 Genomes data in 12 seconds. We also successfully inferred an extended pedigree containing only distant (2nd- and 3rd-degree) relationships representing an aunt, her niece, and her first cousin. Our proposed methods will be implemented in freely distributed software (KING), allowing other investigators to apply the methods directly to analysis of their own sequencing and other high-throughput array data. We expect the relationship inference methods developed here will play an important role in the quality control and analysis of large sets of genetic/genomic data in the coming years.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Family-based rare variant association methods for quantitative traits
-
批准号:8514675
-
项目类别:
-
资助金额:$7.9万
-
财政年份:2012
-
负责人:WEI-MIN CHEN
-
依托单位:
Family-based rare variant association methods for quantitative traits
-
批准号:8355029
-
项目类别:
-
资助金额:$7.9万
-
财政年份:2012
-
负责人:WEI-MIN CHEN
-
依托单位:
海外基金