ModelOMatic: fast and automated model selection between RY, nucleotide, amino acid, and codon substitution models.

ModelOMatic: fast and automated model selection between RY, nucleotide, amino acid, and codon substitution models.
复制标题

ModelOMatic:在 RY、核苷酸、氨基酸和密码子替换模型之间进行快速、自动化的模型选择。

DOI:
10.1093/sysbio/syu062
复制
发表时间:
2015
期刊:
影响因子:
6.5
通讯作者:
Whelan S
Whelan S
中科院分区:
生物学1区
文献类型:
--
作者:
Whelan S

文献摘要

参考文献

被引文献

相似文献

分子生物学是从基因组序列数据推断进化过程和模式的有力工具。统计方法,如最大似然法和贝叶斯推理,现在被确立为首选的推理方法。研究人员用于推理的模型的选择是至关重要的,并且存在基于特定类型的数据(例如核苷酸、氨基酸或密码子)的模型选择的既定方法。现有模型选择方法的一个主要局限性是,它们只能比较作用于单一类型数据的模型。在这里,我们通过引入适配器函数的思想,将聚合模型投影到最初观察到的序列数据上,扩展模型选择以允许描述不同类型数据的模型之间的比较。这些预测在程序ModelOMatic中实现,并用于对来自PANDIT数据库的3722个家族、来自节肢动物基因组数据集的68个基因和来自脊椎动物基因组数据集的248个基因进行模型选择。对于PANDIT和节肢动物的数据,我们发现,氨基酸模型被选择为绝大多数的路线,随着越来越少的路线选择密码子和核苷酸模型,没有家庭选择基于RY的模型。相比之下,几乎所有的脊椎动物数据集的比对选择基于密码子的模型。序列的差异,序列的数量和选择作用于蛋白质序列的程度可能有助于解释模型选择中的这种变化。我们的ModelOMatic程序速度很快,来自PANDIT的大多数家族需要不到150秒才能完成,因此应该很容易纳入现有的系统发育管道。ModelOMatic可在https://code.google.com/p/modelomatic/上获得。
Molecular phylogenetics is a powerful tool for inferring both the process and pattern of evolution from genomic sequence data. Statistical approaches, such as maximum likelihood and Bayesian inference, are now established as the preferred methods of inference. The choice of models that a researcher uses for inference is of critical importance, and there are established methods for model selection conditioned on a particular type of data, such as nucleotides, amino acids, or codons. A major limitation of existing model selection approaches is that they can only compare models acting upon a single type of data. Here, we extend model selection to allow comparisons between models describing different types of data by introducing the idea of adapter functions, which project aggregated models onto the originally observed sequence data. These projections are implemented in the program ModelOMatic and used to perform model selection on 3722 families from the PANDIT database, 68 genes from an arthropod phylogenomic data set, and 248 genes from a vertebrate phylogenomic data set. For the PANDIT and arthropod data, we find that amino acid models are selected for the overwhelming majority of alignments; with progressively smaller numbers of alignments selecting codon and nucleotide models, and no families selecting RY-based models. In contrast, nearly all alignments from the vertebrate data set select codon-based models. The sequence divergence, the number of sequences, and the degree of selection acting upon the protein sequences may contribute to explaining this variation in model selection. Our ModelOMatic program is fast, with most families from PANDIT taking fewer than 150 s to complete, and should therefore be easily incorporated into existing phylogenetic pipelines. ModelOMatic is available at https://code.google.com/p/modelomatic/.
DOI: 10.1093/bioinformatics/btg188
发表时间: 2003-08-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Whelan, S;de Bakker, PIW;Goldman, N
通讯作者: Goldman, N
DOI: 10.1109/tac.1974.1100705
发表时间: 1974-01-01
影响因子: 6.8
作者:
AKAIKE, H
通讯作者: AKAIKE, H
DOI: 10.1093/molbev/msg184
发表时间: 2003-10-01
影响因子: 10.7
作者:
Robinson, DM;Jones, DT;Thorne, JL
通讯作者: Thorne, JL
2. 生物学与进化
DOI: 10.1163/ej.9789004162259.i-352.14
发表时间: 2010
影响因子: 10.7
作者:
S. Subbotin;M. Mundo‐Ocampo;J. Baldwin
通讯作者: J. Baldwin
DOI: 10.1186/1471-2148-11-146
发表时间: 2011-05-27
影响因子: 3.4
作者:
Letsch HO;Kjer KM
通讯作者: Kjer KM