Fregene: Simulation of realistic sequence-level data in populations and ascertained samples

Fregene: Simulation of realistic sequence-level data in populations and ascertained samples
复制标题

DOI:
10.1186/1471-2105-9-364
复制
发表时间:
2008-09-08
期刊:
影响因子:
3
通讯作者:
Balding, David J.
Balding, David J.
中科院分区:
生物学4区
文献类型:
--
作者:
Chadeau-Hyam, Marc;Hoggart, Clive J.;Balding, David J.

文献摘要

被引文献

相似文献

背景:FREGENE模拟大群体中大基因组区域的序列水平数据。因为,与合并模拟器不同,它通过时间向前工作,它允许同时模拟选择,人口统计和重组的复杂场景。在FREGENE中实现了对选择中的位点的详细跟踪,并提供了测试理论预测和获得对选择机制的新见解的机会。我们在这里描述的主要功能的两个FREGENE和SAMPLE,一个同伴程序,可以复制关联study datasets.Results:我们报告了详细的分析,我们已经公开的六个大型模拟数据集。模拟了三种人口统计情景:一种是随机的,一种是通过迁移进行子结构化的,还有一种模拟全球主要人群遗传变异主要特征的复杂情景。对于每一种情况下,有一个中性的模拟,一个复杂的模式selection.Conclusion:FREGENE和模拟的数据集将是有价值的选择,人口和人口遗传参数,以及关联研究的有效性模型的有效性进行评估。它的主要优点是建模的灵活性和计算效率。它是开放源代码和面向对象的。因此,它可以定制和模型的范围扩大。
Background: FREGENE simulates sequence-level data over large genomic regions in large populations. Because, unlike coalescent simulators, it works forwards through time, it allows complex scenarios of selection, demography, and recombination to be modelled simultaneously. Detailed tracking of sites under selection is implemented in FREGENE and provides the opportunity to test theoretical predictions and gain new insights into mechanisms of selection. We describe here main functionalities of both FREGENE and SAMPLE, a companion program that can replicate association study datasets.Results: We report detailed analyses of six large simulated datasets that we have made publicly available. Three demographic scenarios are modelled: one panmictic, one substructured with migration, and one complex scenario that mimics the principle features of genetic variation in major worldwide human populations. For each scenario there is one neutral simulation, and one with a complex pattern of selection.Conclusion: FREGENE and the simulated datasets will be valuable for assessing the validity of models for selection, demography and population genetic parameters, as well as the efficacy of association studies. Its principle advantages are modelling flexibility and computational efficiency. It is open source and object-oriented. As such, it can be customised and the range of models extended.