Calibrating a coalescent simulation of human genome sequence variation

Calibrating a coalescent simulation of human genome sequence variation
复制标题

DOI:
10.1101/gr.3709305
复制
发表时间:
2005-11-01
期刊:
影响因子:
7
通讯作者:
Altshuler, D
Altshuler, D
中科院分区:
生物学1区
文献类型:
--
作者:
Schaffner, SF;Foo, C;Altshuler, D

文献摘要

被引文献

相似文献

群体遗传学模型在人类遗传学研究中起着重要作用,它将有关序列变异的实证观察与关于潜在历史和生物学原因的假设联系起来。更具体地说,这些模型被用于将序列变异、连锁不平衡(LD)和选择的实证测量值与“零”分布下的预期值进行比较。在缺乏有关人类人口统计学历史以及突变和重组率变化的详细信息的情况下,模拟必然使用了任意的模型,通常是简单的模型。随着大量实证数据集的出现,现在有可能利用全基因组数据对群体遗传学模型进行校准,这首次使得能够生成在多种特征上与实证数据一致的数据。我们在此介绍了第一个这样经过校准的模型,并表明尽管它仍然是任意的,但它成功地生成了(针对三个群体的)模拟数据,这些数据在等位基因频率、连锁不平衡和群体分化方面与实证数据非常相似。对于所提出的历史和重组模型的准确性并没有做出断言,但它生成符合实际数据的能力满足了遗传学家长期以来的需求。我们预计这个模型(其软件是公开可用的)以及其他类似的模型将在人类遗传学的实证研究中有众多应用。
Population genetic models play an important role in human genetic research, connecting empirical observations about sequence variation with hypotheses about underlying historical and biological causes. More specifically, models are used to compare empirical measures of sequence variation, linkage disequilibrium (LD), and selection to expectations under a "null" distribution. In the absence of detailed information about human demographic history, and about variation in mutation and recombination rates, simulations have of necessity used arbitrary models, usually simple ones. With the advent of large empirical data sets, it is now possible to calibrate Population genetic models with genome-wide data, permitting for the first time the generation of data that are consistent with empirical data across a wide range of characteristics. We present here the first such calibrated model and show that, while still arbitrary, it successfully generates simulated data (for three populations) that closely resemble empirical data in allele frequency, linkage disequilibrium, and population differentiation. No assertion is made about the accuracy of the proposed historical and recombination model, but its ability to generate realistic data meets a long-standing need among geneticists. We anticipate that this model, for which software is Publicly available, and others like it will have numerous applications in empirical studies of human genetics.