PyRAD: assembly of de novo RADseq loci for phylogenetic analyses

PyRAD: assembly of de novo RADseq loci for phylogenetic analyses
复制标题

DOI:
10.1093/bioinformatics/btu121
复制
发表时间:
2014-07-01
期刊:
影响因子:
5.8
通讯作者:
Eaton, Deren A. R.
Eaton, Deren A. R.
中科院分区:
生物学3区
文献类型:
--
作者:
Eaton, Deren A. R.

文献摘要

被引文献

相似文献

动机:限制位点相关的基因组标记是一个强大的工具,调查在人口水平上的进化问题,但在更深层次的系统发育尺度,其中较少的orthopathic基因座通常回收不同的类群在其效用有限。虽然这种限制部分源于对限制性识别位点的突变,这些突变破坏了数据生成,但数据丢失的另一个来源来自生物信息学分析期间未能识别同源性。允许较低的相似性阈值和包括indel变异的聚类方法将在系统发育尺度上组装RADseq基因座时表现得更好。结果:PyRAD是一个组装从头RADseq基因座的管道,目的是优化系统发育数据集的覆盖率。它使用了一个包装器围绕一个聚类算法,这允许样本内和样本之间的插入缺失变异,以及读段之间的不完全重叠(例如,G.双端)。在这里,我比较了PyRAD和Stacks程序在分析包含indel变异的模拟RADseq数据集时的性能。插入缺失破坏了Stacks中的同源基因座的聚类,但在PyRAD中没有,使得后者在不同的分类群中恢复了更多的共享基因座。我通过对经验RADseq数据集的重新分析表明,插入缺失是此类数据的一个共同特征,即使在浅系统发育尺度上也是如此。PyRAD使用并行处理以及可选的分层聚类方法,这使其能够快速组装具有数百个样本个体的系统发育数据集。可用性:软件是用Python编写的,可在http://www.德雷纳顿公司简介
Motivation: Restriction-site-associated genomic markers are a powerful tool for investigating evolutionary questions at the population level, but are limited in their utility at deeper phylogenetic scales where fewer orthologous loci are typically recovered across disparate taxa. While this limitation stems in part from mutations to restriction recognition sites that disrupt data generation, an additional source of data loss comes from the failure to identify homology during bioinformatic analyses. Clustering methods that allow for lower similarity thresholds and the inclusion of indel variation will perform better at assembling RADseq loci at the phylogenetic scale. Results: PyRAD is a pipeline to assemble de novo RADseq loci with the aim of optimizing coverage across phylogenetic datasets. It uses a wrapper around an alignment-clustering algorithm, which allows for indel variation within and between samples, as well as for incomplete overlap among reads (e. g. paired-end). Here I compare PyRAD with the program Stacks in their performance analyzing a simulated RADseq dataset that includes indel variation. Indels disrupt clustering of homologous loci in Stacks but not in PyRAD, such that the latter recovers more shared loci across disparate taxa. I show through reanalysis of an empirical RADseq dataset that indels are a common feature of such data, even at shallow phylogenetic scales. PyRAD uses parallel processing as well as an optional hierarchical clustering method, which allows it to rapidly assemble phylogenetic datasets with hundreds of sampled individuals. Availability: Software is written in Python and freely available at http://www. dereneaton. com/software/