PyEvolve: a toolkit for statistical modelling of molecular evolution.

PyEvolve: a toolkit for statistical modelling of molecular evolution.
复制标题

DOI:
10.1186/1471-2105-5-1
复制
发表时间:
2004-01-05
期刊:
影响因子:
3
通讯作者:
Huttley GA
Huttley GA
中科院分区:
生物学4区
文献类型:
--
作者:
Butterfield A;Vedagiri V;Lang E;Lawrence C;Wakefield MJ;Isaev A;Huttley GA

文献摘要

参考文献

被引文献

相似文献

检验变异的分布已被证明是一种极其有益的技术,有助于识别具有生物学意义的序列。然而,该领域的大多数方法只评估序列的保守部分,而忽略了序列差异的生物学意义。基于分子进化领域的一套复杂的基于似然的统计模型为从序列变异的全分布中提取信息提供了基础。基于系统发育的最大似然计算可以应用于不同问题的数量是广泛的。可以执行可能性计算的可用软件包缺乏灵活性和可扩展性,或者采用容易出错的方法进行模型参数化。在这里,我们描述了PyEvolve的实现,这是一个用于应用现有和开发新的分子进化统计方法的工具包。提出了PyEvolve的对象体系结构和设计模式,其中包括一个自适应的多级并行化模式。定义新方法的方法通过实现一种新的二核苷酸替代模型来说明,该模型包括甲基化CpG突变的参数,需要8行标准Python代码来定义。使用二核苷酸或密码子替代模型对来自20种哺乳动物或10个物种子集的BRCA1序列进行比对。据记录,与串行相比,并行的性能提升高达5倍。与领先的替代软件相比,PyEvolve在具有大数据集的参数丰富模型上表现出更好的实际性能,将优化所需的时间从10天减少到6小时。PyEvolve提供了灵活的功能,既可以用于分子进化的统计建模,也可以用于开发该领域的新方法。该工具包可以交互式使用,也可以通过编写和执行脚本来使用。该工具包使用有效的过程来指定统计模型的参数化,并实现了许多优化,使高度参数丰富的似然函数可以在数小时内在多cpu硬件上解决。PyEvolve可以很容易地适应不断变化的计算需求和硬件配置,以最大限度地提高性能。PyEvolve根据GPL发布,可以从。
Examining the distribution of variation has proven an extremely profitable technique in the effort to identify sequences of biological significance. Most approaches in the field, however, evaluate only the conserved portions of sequences – ignoring the biological significance of sequence differences. A suite of sophisticated likelihood based statistical models from the field of molecular evolution provides the basis for extracting the information from the full distribution of sequence variation. The number of different problems to which phylogeny-based maximum likelihood calculations can be applied is extensive. Available software packages that can perform likelihood calculations suffer from a lack of flexibility and scalability, or employ error-prone approaches to model parameterisation. Here we describe the implementation of PyEvolve, a toolkit for the application of existing, and development of new, statistical methods for molecular evolution. We present the object architecture and design schema of PyEvolve, which includes an adaptable multi-level parallelisation schema. The approach for defining new methods is illustrated by implementing a novel dinucleotide model of substitution that includes a parameter for mutation of methylated CpG's, which required 8 lines of standard Python code to define. Benchmarking was performed using either a dinucleotide or codon substitution model applied to an alignment of BRCA1 sequences from 20 mammals, or a 10 species subset. Up to five-fold parallel performance gains over serial were recorded. Compared to leading alternative software, PyEvolve exhibited significantly better real world performance for parameter rich models with a large data set, reducing the time required for optimisation from ~10 days to ~6 hours. PyEvolve provides flexible functionality that can be used either for statistical modelling of molecular evolution, or the development of new methods in the field. The toolkit can be used interactively or by writing and executing scripts. The toolkit uses efficient processes for specifying the parameterisation of statistical models, and implements numerous optimisations that make highly parameter rich likelihood functions solvable within hours on multi-cpu hardware. PyEvolve can be readily adapted in response to changing computational demands and hardware configurations to maximise performance. PyEvolve is released under the GPL and can be downloaded from .
DOI: 10.1093/bioinformatics/14.2.219
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
McGuire, G;Wright, F
通讯作者: Wright, F
DOI: 10.1093/nar/22.22.4673
发表时间: 1994-11-11
影响因子: 14.9
作者:
THOMPSON, JD;HIGGINS, DG;GIBSON, TJ
通讯作者: GIBSON, TJ
DOI: 10.1038/35020557
发表时间: 2000-08-10
期刊: NATURE
影响因子: 64.8
作者:
Bohossian, HB;Skaletsky, H;Page, DC
通讯作者: Page, DC
DOI: 10.2307/2532163
发表时间: 1991-06-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
HALL, P;WILSON, SR
通讯作者: WILSON, SR
DOI: 10.1038/385151a0
发表时间: 1997-01-09
期刊: NATURE
影响因子: 64.8
作者:
Messler, W;Stewart, CB
通讯作者: Stewart, CB