Genome-wide detection and characterization of positive selection in human populations

Genome-wide detection and characterization of positive selection in human populations
复制标题

DOI:
10.3410/f.1094912.549862
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
D. Ballinger;D. Hinds;Andrew Boudreau;Suzanne M. Leal;S. Pasternak;David A. Wheeler;Thomas D. Willis-Thomas-D.
D. Ballinger;D. Hinds;Andrew Boudreau;Suzanne M. Leal;S. Pasternak;David A. Wheeler;Thomas D. Willis-Thomas-D.
中科院分区:
其他
文献类型:
--
作者:
D. Ballinger;D. Hinds;Andrew Boudreau;Suzanne M. Leal;S. Pasternak;David A. Wheeler;Thomas D. Willis-Thomas-D.

文献摘要

被引文献

相似文献

随着密集的人类遗传变异图谱的出现,现在有可能在整个人类基因组中检测到积极的自然选择。在这里,我们报告了来自国际HapMap项目第二阶段(HapMap2)的300多万个多态性的分析。我们使用了‘长程单倍型’方法,这种方法是用来识别在最近经历过选择的种群中分离的等位基因,我们还开发了基于跨种群比较的新方法,以发现种群内几乎固定的等位基因。该分析揭示了300多个强有力的候选地区。聚焦于最强的22个区域,我们开发了一个启发式算法来仔细检查这些区域,以确定候选的选择目标。在互补性分析中,我们确定了26个非同义、编码的单核苷酸多态,显示了正选择的区域证据。对这些候选基因的研究突出了三个案例,其中共同生物过程中的两个基因显然在同一人群中经历了正向选择:西非的Large和DMD都与拉萨病毒的感染有关;欧洲的SLC24A5和SLC45A2都与皮肤色素沉着有关;以及亚洲的EDAR和EDA2R都与毛囊的发育有关。越来越多的关于遗传变异的信息,加上新的分析方法,使探索人类种群最近的进化史成为可能。国际单倍型图谱的第一阶段,包括100万个单核苷酸多态(SNPs),允许对人类的自然选择进行初步检查。现在,随着第二阶段图谱(HapMap2)的发表,来自三大洲的420条染色体(120条欧洲人(CEU)、120条非洲人(YRI)和180条亚洲人(日本)和中国(JPT 1 CHB))的300多万个SNPs已经被分型。在我们对HapMap2的分析中,我们首先实现了两个广泛使用的测试,通过寻找在异常长的单倍型上进行的共同等位基因来检测最近的正选择。长程单倍型(LRH)和综合单倍型评分(IHS)测试依赖的原理是,在正选择下,等位基因可能会足够快地上升到高频,以至于与附近多态的远距离关联-长程单倍型-将没有时间通过重组来消除。这些测试通过将长单倍型与同一基因座上的其他等位基因进行比较来控制重组率的局部变异。结果,随着选定的等位基因接近固定(100%频率),它们失去了力量,因为在种群中几乎没有可供选择的等位基因(补充图2和补充表1-2)。接下来,我们开发、评估和应用了一种新的测试,跨群体扩展单倍型同源性(XP-EHH),以检测选择的等位基因在一个群体中接近或实现固定,但在整个人类群体中保持多态的选择性扫描(方法,以及补充图2和补充表3-6)。最近还描述了相关方法。我们使用这三种方法对最近的正选择进行分析,揭示了300多个候选区(补充图3和补充表7),其中22个高于阈值,使得在10 GB的模拟中性进化序列(方法)中没有发现类似的事件。我们重点研究了这22个最强的信号(表1),其中包括两个公认的案例,SLC24A5和LCT,以及其他20个信号强度相似的区域。挑战是筛选候选区域的遗传变异,以确定作为选择目标的变异。我们的候选区域很大(平均长度为815kb;最大长度为3.5Mb),并且通常包含多个基因(中位数为4;最大为15)。一个典型的地区拥有400-4,000个常见的SNP(次要等位基因频率为0.5%),其中大约四分之三在当前的SNP数据库中出现,一半作为HapMap2的一部分进行了基因分型(补充表8)。我们制定了三个标准来帮助突出潜在的选择目标(补充图1):(1)通过我们的测试可检测到的选定等位基因很可能是派生的(新出现的),因为长期的单倍型测试几乎没有能力检测现有(先前存在的)变异的选择;因此我们专注于派生的等位基因,通过与灵长类外群的比较来识别;(2)选定的等位基因很可能在不同群体之间高度分化,因为最近的选择可能是一种局部环境适应;因此我们只寻找在被选择的群体(S)中常见的等位基因;(3)选定的等位基因必须具有生物效应。因此,在现有知识的基础上,我们重点研究了非同义编码的SNPs和进化保守序列中的SNPs。这些标准是启发式的,而不是绝对的要求。有些选择的目标可能不能满足他们的要求,有些则不会出现在目前的SNP数据库中。尽管如此,由于这些人群中50%的常见SNP是在HapMap2中进行基因分型的,寻找因果变异是及时的。我们将标准应用于包含SLC24A5和LCT的区域,每个区域都已经有很强的候选基因、突变和性状。在SLC24A5,600kb区域包含914个基因分型
With the advent of dense maps of human genetic variation, it is now possible to detect positive natural selection across the human genome. Here we report an analysis of over 3 million polymorphisms from the International HapMap Project Phase 2 (HapMap2). We used ‘long-range haplotype’ methods, which were developed to identify alleles segregating in a population that have undergone recent selection, and we also developed new methods that are based on cross-population comparisons to discover alleles that have swept to near-fixation within a population. The analysis reveals more than 300 strong candidate regions. Focusing on the strongest 22 regions, we develop a heuristic for scrutinizing these regions to identify candidate targets of selection. In a complementary analysis, we identify 26 nonsynonymous, coding, single nucleotide polymorphisms showing regional evidence of positive selection. Examination of these candidates highlights three cases in which two genes in a common biological process have apparently undergone positive selection in the same population: LARGE and DMD, both related to infection by the Lassa virus, in West Africa; SLC24A5 and SLC45A2, both involved in skin pigmentation, in Europe; and EDAR and EDA2R, both involved in development of hair follicles, in Asia. An increasing amount of information about genetic variation, together with new analytical methods, is making it possible to explore the recent evolutionary history of the human population. The first phase of the International Haplotype Map, including ,1 million single nucleotide polymorphisms (SNPs), allowed preliminary examination of natural selection in humans. Now, with the publication of the Phase 2 map (HapMap2) in a companion paper, over 3 million SNPs have been genotyped in 420 chromosomes from three continents (120 European (CEU), 120 African (YRI) and 180 Asian from Japan and China (JPT 1 CHB)). In our analysis of HapMap2, we first implemented two widely used tests that detect recent positive selection by finding common alleles carried on unusually long haplotypes. The two, the Long-Range Haplotype (LRH) and the integrated Haplotype Score (iHS) tests, rely on the principle that, under positive selection, an allele may rise to high frequency rapidly enough that long-range association with nearby polymorphisms—the long-range haplotype—will not have time to be eliminated by recombination. These tests control for local variation in recombination rates by comparing long haplotypes to other alleles at the same locus. As a result, they lose power as selected alleles approach fixation (100% frequency), because there are then few alternative alleles in the population (Supplementary Fig. 2 and Supplementary Tables 1–2). We next developed, evaluated and applied a new test, Cross Population Extended Haplotype Homozogysity (XP-EHH), to detect selective sweeps in which the selected allele has approached or achieved fixation in one population but remains polymorphic in the human population as a whole (Methods, and Supplementary Fig. 2 and Supplementary Tables 3–6). Related methods have recently also been described. Our analysis of recent positive selection, using the three methods, reveals more than 300 candidate regions(Supplementary Fig. 3 and Supplementary Table 7), 22 of which are above a threshold such that no similar events were found in 10 Gb of simulated neutrally evolving sequence (Methods). We focused on these 22 strongest signals (Table 1), which include two well-established cases, SLC24A5 and LCT, and 20 other regions with signals of similar strength. The challenge is to sift through genetic variation in the candidate regions to identify the variants that were the targets of selection. Our candidate regions are large (mean length, 815 kb; maximum length, 3.5 Mb) and often contain multiple genes (median, 4; maximum, 15). A typical region harbours ,400–4,000 common SNPs (minor allele frequency .5%), of which roughly three-quarters are represented in current SNP databases and half were genotyped as part of HapMap2 (Supplementary Table 8). We developed three criteria to help highlight potential targets of selection (Supplementary Fig. 1): (1) selected alleles detectable by our tests are likely to be derived (newly arisen), because long-haplotype tests have little power to detect selection on standing (pre-existing) variation; we therefore focused on derived alleles, as identified by comparison to primate outgroups; (2) selected alleles are likely to be highly differentiated between populations, because recent selection is probably a local environmental adaptation; we thus looked for alleles common in only the population(s) under selection; (3) selected alleles must have biological effects. On the basis of current knowledge, we therefore focused on non-synonymous coding SNPs and SNPs in evolutionarily conserved sequences. These criteria are intended as heuristics, not absolute requirements. Some targets of selection may not satisfy them, and some will not be in current SNP databases. Nonetheless, with ,50% of common SNPs in these populations genotyped in HapMap2, a search for causal variants is timely. We applied the criteria to the regions containing SLC24A5 and LCT, each of which already has a strong candidate gene, mutation and trait. At SLC24A5, the 600 kb region contains 914 genotyped