AMAS: a fast tool for alignment manipulation and computing of summary statistics.

AMAS: a fast tool for alignment manipulation and computing of summary statistics.
复制标题

DOI:
10.7717/peerj.1660
复制
发表时间:
2016
期刊:
影响因子:
2.7
通讯作者:
Borowiec ML
Borowiec ML
中科院分区:
生物学3区
文献类型:
--
作者:
Borowiec ML

文献摘要

被引文献

相似文献

近几年来,遗传学研究中所使用的数据量呈爆炸式增长,许多遗传学推断涉及数百甚至数千个基因座和许多分类群。这些现代基因组学研究除了对基因子集或串联序列进行多重分析外,还需要对每个基因座进行单独分析。需要用于处理和计算数千个单基因座或大串联比对的计算有效的工具。在这里,我介绍了AMAS(对齐操作和摘要),这是一个可以作为独立命令行实用程序或Python包使用的工具。AMAS工作于氨基酸和核苷酸比对,并将序列操作的能力与计算基本统计的功能相结合。操作功能包括流行格式之间的转换,连接,提取网站和分裂根据预定义的分区方案,创建复制数据集,并删除分类。计算的统计数据包括分类群的数量、比对长度、基质细胞的总计数、未确定字符的总数、缺失数据的百分比、AT和GC含量(用于DNA比对)、可变位点的计数和比例、简约信息位点的计数和比例以及与核苷酸或氨基酸字母表相关的所有字符的计数。AMAS特别适合于具有数百个分类群和数千个基因座的非常大的比对。它是计算效率高,利用并行处理,并执行更好的串联比其他流行的工具。AMAS是一个Python 3程序,它完全依赖于Python的核心模块,不需要额外的依赖项。AMAS源代码和手册可以在GNU通用公共许可证下从下载。
The amount of data used in phylogenetics has grown explosively in the recent years and many phylogenies are inferred with hundreds or even thousands of loci and many taxa. These modern phylogenomic studies often entail separate analyses of each of the loci in addition to multiple analyses of subsets of genes or concatenated sequences. Computationally efficient tools for handling and computing properties of thousands of single-locus or large concatenated alignments are needed. Here I present AMAS (Alignment Manipulation And Summary), a tool that can be used either as a stand-alone command-line utility or as a Python package. AMAS works on amino acid and nucleotide alignments and combines capabilities of sequence manipulation with a function that calculates basic statistics. The manipulation functions include conversions among popular formats, concatenation, extracting sites and splitting according to a pre-defined partitioning scheme, creation of replicate data sets, and removal of taxa. The statistics calculated include the number of taxa, alignment length, total count of matrix cells, overall number of undetermined characters, percent of missing data, AT and GC contents (for DNA alignments), count and proportion of variable sites, count and proportion of parsimony informative sites, and counts of all characters relevant for a nucleotide or amino acid alphabet. AMAS is particularly suitable for very large alignments with hundreds of taxa and thousands of loci. It is computationally efficient, utilizes parallel processing, and performs better at concatenation than other popular tools. AMAS is a Python 3 program that relies solely on Python’s core modules and needs no additional dependencies. AMAS source code and manual can be downloaded from under GNU General Public License.