Recommendations for utilizing and reporting population genetic analyses: the reproducibility of genetic clustering using the program STRUCTURE

Recommendations for utilizing and reporting population genetic analyses: the reproducibility of genetic clustering using the program STRUCTURE
复制标题

DOI:
10.1111/j.1365-294x.2012.05754.x
复制
发表时间:
2012-10-01
期刊:
影响因子:
4.9
通讯作者:
Vines, Timothy H.
Vines, Timothy H.
中科院分区:
生物学1区
文献类型:
--
作者:
Gilbert, Kimberly J.;Andrew, Rose L.;Vines, Timothy H.

文献摘要

被引文献

相似文献

可重复性是科学研究结果和结论的基准,但对科学结果可重复性的系统研究却出奇地少。此外,许多现代统计方法利用“随机游走”模型拟合过程,这些方法的输出具有内在的随机性。这些统计程序与当前数据存档和方法报告标准的结合是否允许复制作者的结果?为了验证这一点,我们使用软件包STRUCTURE重新分析了从论文中收集的数据集,以识别遗传相似的个体群。我们发现,再现结构的结果可以是困难的,尽管程序的直接要求。我们的研究结果表明,30%的分析无法重现相同数量的人口集群。为了改善这一点,我们提出了建议,为今后使用的软件和报告结构分析和结果发表的作品。
Reproducibility is the benchmark for results and conclusions drawn from scientific studies, but systematic studies on the reproducibility of scientific results are surprisingly rare. Moreover, many modern statistical methods make use of 'random walk' model fitting procedures, and these are inherently stochastic in their output. Does the combination of these statistical procedures and current standards of data archiving and method reporting permit the reproduction of the authors' results? To test this, we reanalysed data sets gathered from papers using the software package STRUCTURE to identify genetically similar clusters of individuals. We find that reproducing STRUCTURE results can be difficult despite the straightforward requirements of the program. Our results indicate that 30% of analyses were unable to reproduce the same number of population clusters. To improve this, we make recommendations for future use of the software and for reporting STRUCTURE analyses and results in published works.