Benchmarking bacterial genome-wide association study methods using simulated genomes and phenotypes

Benchmarking bacterial genome-wide association study methods using simulated genomes and phenotypes
复制标题

DOI:
10.1099/mgen.0.000337
复制
发表时间:
2020-03-01
期刊:
影响因子:
3.9
通讯作者:
Shapiro, B. Jesse
Shapiro, B. Jesse
中科院分区:
生物学2区
文献类型:
--
作者:
Saber, Morteza M.;Shapiro, B. Jesse

文献摘要

被引文献

相似文献

全基因组关联研究(GWAS)有可能揭示微生物表型的遗传学,如抗生素耐药性和毒力。利用不断增长的细菌序列数据,微生物GWAS方法旨在识别因果遗传变异,同时忽略虚假关联。细菌克隆繁殖,导致强大的群体结构和全基因组连锁,使得将真正的“命中”(即导致表型的突变)与非因果连锁突变分开具有挑战性。GWAS方法试图以不同的方式校正种群结构,但它们的性能尚未在一系列进化情景下进行系统和全面的评估。在这里,我们开发了一种细菌GWAS模拟器(BacGWASim),以产生具有不同突变率、重组和其他进化参数的细菌基因组,沿着一个感兴趣表型的因果突变子集。我们评估了三种广泛使用的单位点GWAS方法(基于聚类的,降维和线性混合模型,在plink,pyseer和gemma中实现)和一种相对较新的多位点模型在pyseer中实现的性能(召回率和精度),在一系列模拟样本量,重组率和因果突变效应大小范围内。正如预期的那样,所有方法在样本量和效应量较大时表现更好。根据参数的选择,聚类和降维方法校正种群结构的性能变化很大。值得注意的是,多位点弹性网(lasso)方法始终是性能最高的方法之一,并且在检测具有低和高效应量的因果变异方面具有最高的功效。大多数方法达到了良好的性能水平(召回>0.75),用于识别强效应大小[对数比值比(OR)>= 2]的因果突变,样本量为2000个基因组。然而,只有弹性网络达到了合理的性能水平(召回率=0.35),用于检测较小样本中具有较弱效果(log OR类似于1)的标记。相对于单位点模型,弹性网络在控制全基因组连锁方面也表现出上级的精确度和召回率。然而,所有的方法表现相对较差的高度克隆(低重组)的基因组,这表明在方法开发的改进空间。这些发现显示了多位点模型改善细菌GWAS性能的潜力。BacGWASim代码和模拟数据是公开的,可以对新方法进行进一步的比较和基准测试。
Genome-wide association studies (GWASs) have the potential to reveal the genetics of microbial phenotypes such as antibiotic resistance and virulence. Capitalizing on the growing wealth of bacterial sequence data, microbial GWAS methods aim to identify causal genetic variants while ignoring spurious associations. Bacteria reproduce clonally, leading to strong population structure and genome-wide linkage, making it challenging to separate true 'hits' (i.e. mutations that cause a phenotype) from non-causal linked mutations. GWAS methods attempt to correct for population structure in different ways, but their performance has not yet been systematically and comprehensively evaluated under a range of evolutionary scenarios. Here, we developed a bacterial GWAS simulator (BacGWASim) to generate bacterial genomes with varying rates of mutation, recombination and other evolutionary parameters, along with a subset of causal mutations underlying a phenotype of interest. We assessed the performance (recall and precision) of three widely used single-locus GWAS approaches (cluster-based, dimensionality-reduction and linear mixed models, implemented in plink, pyseer and gemma) and one relatively new multi-locus model implemented in pyseer, across a range of simulated sample sizes, recombination rates and causal mutation effect sizes. As expected, all methods performed better with larger sample sizes and effect sizes. The performance of clustering and dimensionality reduction approaches to correct for population structure were considerably variable according to the choice of parameters. Notably, the multi-locus elastic net (lasso) approach was consistently amongst the highest-performing methods, and had the highest power in detecting causal variants with both low and high effect sizes. Most methods reached the level of good performance (recall >0.75) for identifying causal mutations of strong effect size [log odds ratio (OR) >= 2] with a sample size of 2000 genomes. However, only elastic nets reached the level of reasonable performance (recall=0.35) for detecting markers with weaker effects (log OR similar to 1) in smaller samples. Elastic nets also showed superior precision and recall in controlling for genome-wide linkage, relative to single-locus models. However, all methods performed relatively poorly on highly clonal (low-recombining) genomes, suggesting room for improvement in method development. These findings show the potential for multi-locus models to improve bacterial GWAS performance. BacGWASim code and simulated data are publicly available to enable further comparisons and benchmarking of new methods.