Linkage disequilibrium and inference of ancestral recombination in 538 single-nucleotide polymorphism clusters across the human genome

Linkage disequilibrium and inference of ancestral recombination in 538 single-nucleotide polymorphism clusters across the human genome
复制标题

DOI:
10.1086/377138
复制
发表时间:
2003-08-01
影响因子:
9.8
通讯作者:
Lai, E
Lai, E
中科院分区:
生物学1区
文献类型:
--
作者:
Clark, AG;Nielsen, R;Lai, E

文献摘要

被引文献

相似文献

利用连锁不平衡(LD)在人类中进行精细定位的前景已经引起了相当大的关注,在验证一组用于连锁分析的单核苷酸多态(SNPs)的过程中,产生了一组538个簇中4833个SNPs的数据,提供了LD在整个基因组中的局部属性的丰富图景。LD估计可能有偏差,这取决于最初确定SNPs的方法,当随后在较大的人口样本中对小的异质小组中确定的SNPs进行分型时,就会出现确定偏差的特殊问题。理解和纠正确定偏差对于对整个人类基因组中LD的情况进行有用的定量评估是至关重要的。基因组上群体重组率的异质性,Rho=4nR,反映了为了获得最佳覆盖,标记密度必须有多大的变化。我们发现,确定校正后的Rho在基因组上的变化超过两个数量级,这意味着我们基因组不同部分的重组历史有很大的差异。(Rho)在帽上的分布是单峰的,我们表明这与在可变复合速率的背景下的大范围的热点混合是相容的。虽然(Rho)Over CAP在三个群体样本中显著相关,但基因组的某些区域在Rho中表现出群体特有的尖峰或低谷,这些尖峰或低谷太大,无法通过采样来解释。这一结果与当地基因组区域谱系深度的差异是一致的,这一发现直接关系到LD图谱的设计和应用,也与美国国立卫生研究院的HapMap项目有关。
The prospect of using linkage disequilibrium (LD) for fine-scale mapping in humans has attracted considerable attention, and, during the validation of a set of single-nucleotide polymorphisms ( SNPs) for linkage analysis, a set of data for 4,833 SNPs in 538 clusters was produced that provides a rich picture of local attributes of LD across the genome. LD estimates may be biased depending on the means by which SNPs are first identified, and a particular problem of ascertainment bias arises when SNPs identified in small heterogeneous panels are subsequently typed in larger population samples. Understanding and correcting ascertainment bias is essential for a useful quantitative assessment of the landscape of LD across the human genome. Heterogeneity in the population recombination rate, rho = 4Nr, along the genome reflects how variable the density of markers will have to be for optimal coverage. We find that ascertainment-corrected rho varies along the genome by more than two orders of magnitude, implying great differences in the recombinational history of different portions of our genome. The distribution of (rho) over cap is unimodal, and we show that this is compatible with a wide range of mixtures of hotspots in a background of variable recombination rate. Although (rho) over cap is significantly correlated across the three population samples, some regions of the genome exhibit population-specific spikes or troughs in rho that are too large to be explained by sampling. This result is consistent with differences in the genealogical depth of local genomic regions, a finding that has direct bearing on the design and utility of LD mapping and on the National Institutes of Health HapMap project.