Population Structure Admixture and Selection across the 1000 Genomes Data Set
Population Structure Admixture and Selection across the 1000 Genomes Data Set
批准号:
8526601
负责人:
Carlos Daniel Bustamante
金额:
$19.63万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-09-09 至 2014-06-30
关键词:
AccountingAdmixtureAfrican AmericanComplexComputing MethodologiesDataData AnalysesData SetDevelopmentDiffusionDiseaseEquilibriumFrequenciesFutureGene FrequencyGeneticGenetic ResearchGenetic VariationGenomeGenomicsHaplotypesHeatingHereditary DiseaseHispanicsHumanInternationalJointsLatinoMapsMedicalMedical GeneticsMethodsMinority GroupsModelingNatural SelectionsPatternPopulationPopulation GeneticsReadingRecording of previous eventsResearchSamplingShapesSignal TransductionSiteSoftware ToolsStructureSurveysTestingVariantbasedesigndisorder riskethnic minority populationgenome sequencinggenome wide association studygenome-widehuman population geneticshuman subjectnovelnovel strategiessoftware development
中文摘要
描述(由申请人提供):1000个基因组计划(TGP)在回答人类群体遗传学的基本问题和塑造医学基因组研究的未来设计方面具有巨大的潜力。实现这一潜力的关键是开发高效、鲁棒和强大的计算方法,用于分析项目产生的大量数据。在这里,我们提出了新的方法来表征人口结构,分析模式的混合物,并在2,000个样本的三峡工程选择本地化的签名。我们的项目有三个主要目标。首先,我们将以三峡工程为基础,构建详细的人类人口历史模型。为了实现这一目标,我们开发了用于分析所有被调查人群中罕见和常见SNP、拷贝数变异(CNV)和单倍型的联合等位基因频谱的方法。拥有完整的序列数据将使这些方法在对最近的过去进行推断时显着更好,其中频谱的失真对于测试与罕见变异的关联特别重要。其次,我们将在四个西班牙裔/拉丁裔和三个非洲裔美国人TGP样本的人口结构和混合的特点模式。三峡工程为促进这些重要的和未被充分研究的少数民族群体的人口和医学基因组学研究提供了巨大的机会。我们将开发新的统计基因组学方法来重建混合种群的遗传历史,并将这些方法应用于三峡工程样本。我们的方法将针对短读段序列数据量身定制,并将利用采样的三重设计。第三,我们将在完整的TGP数据集中检测平衡,净化和积极选择的签名。我们将开发软件工具,以整合自然选择的签名,基于一种新的方法,使用数值方法来拟合多维站点频谱的扩散近似。这种方法允许识别由积极选择、平衡选择或消极选择引起的扭曲。该方法特别适用于低覆盖度短读段序列数据。这些推断将与GWAS命中的地图相结合,以加速发现疾病相关的变异。
相关性:医学遗传学研究为揭示复杂疾病的遗传基础提供了一种工具。1000个基因组计划是一项国际努力,旨在对大约2,000个不同人类受试者的基因组进行测序。我们建议分析这些数据,以表征基因组之间的差异,并促进世界各地的医学和人口基因组研究。
英文摘要
DESCRIPTION (provided by applicant): The 1000 Genomes Project (TGP) has tremendous potential to answer fundamental questions in human population genetics and shape the future design of medical genomic studies. Key to realizing this potential is the development of efficient, robust, and powerful computational methods for analysis of the copious amounts of data generated by the project. Here, we propose novel approaches for characterizing population structure, analyzing patterns of admixture, and localizing signatures of selection across the 2,000 samples of the TGP. Our project has three primary aims. First, we will construct detailed models of human demographic history based on the TGP. To accomplish this, we develop approaches for analyzing the joint allele frequency spectrum of rare and common SNPs, copy number variants (CNVs), and haplotypes across all the populations being surveyed. Having full sequence data will render these approaches dramatically better at making inferences about the recent past, where distortions in frequency spectra are particularly important for testing associations with rare variants. Second, we will characterize patterns of population structure and admixture in the four Hispanic/Latino and three African-American TGP samples. The TGP presents a tremendous opportunity for catalyzing population and medical genomics research for these important and understudied ethnic minority groups. We will develop novel statistical genomic approaches for reconstructing the genetic history of admixed populations and apply these methods to the TGP samples. Our methods will be tailored for short-read sequence data and will leverage the trio design of the sampling. Third, we will detect signatures of balancing, purifying, and positive selection in the full TGP data set. We will develop software tools to integrate signatures of natural selection based on a new approach that uses numerical methods to fit a diffusion approximation to the multi-dimensional site frequency spectrum. This approach allows identification of distortions caused by positive, balancing, or negative selection. The method is especially well suited to low coverage short-read sequence data. These inferences will be integrated with the maps of GWAS hits to accelerate discovery of disease-associated variants.
RELEVANCE: Medical genetics research provides a vehicle for uncovering the heritable basis of complex disease. The 1000 Genomes project is an international effort to sequence the genomes of approximately 2,000 diverse human subjects. We propose to analyze these data in order to characterize differences among genomes and catalyze medical and population genomic research throughout the world.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1534/genetics.113.153973
发表时间:
2013-11
期刊:
Genetics
影响因子:
3.3
作者:
[Gazave E, Chang D, Clark AG, Keinan A]
通讯作者:
Keinan A
DOI:
10.1093/gbe/evr076
发表时间:
2011
期刊:
Genome biology and evolution
影响因子:
3.3
作者:
[Waldman YY, Tuller T, Keinan A, Ruppin E]
通讯作者:
Ruppin E
Biorepository of Human iPSCs for Studying Dilated and Hypertrophic Cardiomyopathy
-
批准号:9031800
-
项目类别:
-
资助金额:$186.19万
-
财政年份:2014
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Why We Can't Wait: Conference to Eliminate Health Disparities in Genomics
-
批准号:8785928
-
项目类别:
-
资助金额:$5.0万
-
财政年份:2014
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Methods for high-resolution analysis of genetic effects on gene expression
-
批准号:9270646
-
项目类别:
-
资助金额:$33.01万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Methods for high-resolution analysis of genetic effects on gene expression
-
批准号:8915307
-
项目类别:
-
资助金额:$12.32万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Methods for high-resolution analysis of genetic effects on gene expression
-
批准号:8585947
-
项目类别:
-
资助金额:$57.63万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Clinically Relevant Genome Variation Database
-
批准号:8738706
-
项目类别:
-
资助金额:$235.2万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Why We Cant Wait: Conference to Eliminate Health Disparities in Genomics
-
批准号:8529747
-
项目类别:
-
资助金额:$4.39万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Methods for high-resolution analysis of genetic effects on gene expression
-
批准号:8915306
-
项目类别:
-
资助金额:$14.2万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Methods for high-resolution analysis of genetic effects on gene expression
-
批准号:8894321
-
项目类别:
-
资助金额:$62.29万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Methods for high-resolution analysis of genetic effects on gene expression
-
批准号:8711566
-
项目类别:
-
资助金额:$54.44万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Clinically Relevant Genome Variation Database
-
批准号:8574128
-
项目类别:
-
资助金额:$140.0万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Clinically Relevant Genome Variation Database
-
批准号:9047616
-
项目类别:
-
资助金额:$24.95万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Clinically Relevant Genome Variation Database
-
批准号:9134491
-
项目类别:
-
资助金额:$223.47万
-
财政年份:2013
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Genomic Origins and Admixture in Latinos (GOAL)
-
批准号:8327128
-
项目类别:
-
资助金额:$46.15万
-
财政年份:2011
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Genomic Origins and Admixture in Latinos (GOAL)
-
批准号:8108971
-
项目类别:
-
资助金额:$38.88万
-
财政年份:2011
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Genomic Origins and Admixture in Latinos (GOAL)
-
批准号:8535169
-
项目类别:
-
资助金额:$36.76万
-
财政年份:2011
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Genomic Origins and Admixture in Latinos (GOAL)
-
批准号:8727589
-
项目类别:
-
资助金额:$38.67万
-
财政年份:2011
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Population Structure Admixture and Selection across the 1000 Genomes Data Set
-
批准号:8139948
-
项目类别:
-
资助金额:$43.61万
-
财政年份:2010
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Population Structure Admixture and Selection across the 1000 Genomes Data Set
-
批准号:7881973
-
项目类别:
-
资助金额:$44.19万
-
财政年份:2010
-
负责人:Carlos Daniel Bustamante
-
依托单位:
Population Genetic Inferences from Dense Genotype Data
-
批准号:7921193
-
项目类别:
-
资助金额:$41.93万
-
财政年份:2009
-
负责人:Carlos Daniel Bustamante
-
依托单位:
海外基金