Biological sequence simulation for testing complex evolutionary hypotheses: indel-Seq-Gen version 2.0.

Biological sequence simulation for testing complex evolutionary hypotheses: indel-Seq-Gen version 2.0.
复制标题

DOI:
10.1093/molbev/msp174
复制
发表时间:
2009-11
影响因子:
10.7
通讯作者:
Moriyama EN
Moriyama EN
中科院分区:
生物学1区
文献类型:
--
作者:
Strope CL;Abel K;Scott SD;Moriyama EN

文献摘要

参考文献

被引文献

相似文献

序列模拟是验证生物学假说以及检验各种生物信息学和分子进化方法的重要工具。假设检验依赖于序列模拟方法的代表性能力。简单的假设是可以通过模拟随机的、均匀演化的序列集来检验的。然而,测试复杂的假设,例如,局部相似性,需要在异构模型下模拟序列进化。为此,我们先前引入了indel-Seq-Gen版本1.0(iSGv1.0; indel,插入/缺失)。iSGv1.0允许异质性蛋白质进化和基序保守,以及插入和缺失的限制。尽管取得了这些进展,但对于复杂的假设检验,iSGv1.0和其他当前可用的序列模拟方法都是不够的。indel-Seq-Gen 2.0版(iSGv2.0)旨在模拟高度不同的DNA序列和蛋白质超家族的进化。iSGv2.0在iSGv1.0的基础上进行了改进,增加了谱系特异性进化、使用PROSITE-like正则表达式的基序保守、插入缺失跟踪、序列长度约束以及编码和非编码DNA进化。此外,我们形式化的序列表示用于iSGv2.0和发现的缺陷,在建模中使用的indels在当前的最先进的方法,这偏见的假设涉及indels的模拟结果。我们通过使用一种新颖的离散步进过程来修复iSGv2.0中的此缺陷。最后,我们给出了一个例子模拟的calycin超家族序列,并比较iSGv2.0的性能与iSGv1.0和序列进化的随机模型。
Sequence simulation is an important tool in validating biological hypotheses as well as testing various bioinformatics and molecular evolutionary methods. Hypothesis testing relies on the representational ability of the sequence simulation method. Simple hypotheses are testable through simulation of random, homogeneously evolving sequence sets. However, testing complex hypotheses, for example, local similarities, requires simulation of sequence evolution under heterogeneous models. To this end, we previously introduced indel-Seq-Gen version 1.0 (iSGv1.0; indel, insertion/deletion). iSGv1.0 allowed heterogeneous protein evolution and motif conservation as well as insertion and deletion constraints in subsequences. Despite these advances, for complex hypothesis testing, neither iSGv1.0 nor other currently available sequence simulation methods is sufficient. indel-Seq-Gen version 2.0 (iSGv2.0) aims at simulating evolution of highly divergent DNA sequences and protein superfamilies. iSGv2.0 improves upon iSGv1.0 through the addition of lineage-specific evolution, motif conservation using PROSITE-like regular expressions, indel tracking, subsequence-length constraints, as well as coding and noncoding DNA evolution. Furthermore, we formalize the sequence representation used for iSGv2.0 and uncover a flaw in the modeling of indels used in current state of the art methods, which biases simulation results for hypotheses involving indels. We fix this flaw in iSGv2.0 by using a novel discrete stepping procedure. Finally, we present an example simulation of the calycin-superfamily sequences and compare the performance of iSGv2.0 with iSGv1.0 and random model of sequence evolution.
DOI: 10.1093/bioinformatics/8.3.275
发表时间: 1992-06-01
期刊: COMPUTER APPLICATIONS IN THE BIOSCIENCES
影响因子: --
作者:
JONES, DT;TAYLOR, WR;THORNTON, JM
通讯作者: THORNTON, JM
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1093/molbev/msl195
发表时间: 2007-03-01
影响因子: 10.7
作者:
Strope, Cory L.;Scott, Stephen D.;Moriyama, Etsuko N.
通讯作者: Moriyama, Etsuko N.
DOI: 10.1186/1471-2105-6-66
发表时间: 2005-03-22
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Subramanian, AR;Weyer-Menkhoff, J;Morgenstern, B
通讯作者: Morgenstern, B
DOI: 10.1093/bioinformatics/14.2.157
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stoye, J;Evers, D;Meyer, F
通讯作者: Meyer, F