PAML 5: A friendly and powerful bioinformatics resource for phylogenomics
PAML 5: A friendly and powerful bioinformatics resource for phylogenomics
批准号:
BB/X018571/1
负责人:
Ziheng Yang
金额:
$63.77万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --
中文摘要
PAML(最大似然系统发育分析)是一个生物信息学软件工具,广泛应用于分子进化、分子系统发育、病毒学、生物化学和基因组学等领域。目前,它每年吸引超过1000次引用(自1993年首次发布以来,总共有16000次引用),并拥有完善的用户基础。该软件包的主要优势之一是其丰富的DNA和蛋白质序列进化复杂模型集合,这在系统发育的最大似然和贝叶斯方法中很有用。该软件包可用于比较不同的进化树,推断影响蛋白质编码基因的适应性分子进化,重建已灭绝祖先物种的序列等。尽管PAML是一种广泛使用的生物信息学工具,但它的用户界面很差,学习曲线也很陡峭。它不是并行的,并且在应用于大型数据集时计算效率低下。这些问题阻碍了它被新用户广泛采用,并且意味着现有用户必须忍受计算负担。在这个项目中,我们建议重新设计和重新实现PAML程序中的关键算法,并行化代码,并开发一个R接口。我们将通过开发新的密码子进化突变选择模型来扩展和完善其功能,从而实现更准确的祖先序列重建和更稳健的正选择基因检测。这些改进将大大提高程序包的可用性和计算性能,使其成为生物科学研究界极有价值的生物信息学资源。本项目是低风险、高影响的项目。HPC技术、pthread和MPI已经存在很多年了。在系统发育学中,所有主要的似然和贝叶斯程序都已经成功地并行化了,包括PhyML、RAxML、MrBayes/RevBayes和BEAST。实际上,PAML在这个意义上是唯一糟糕的,因为它没有被并行化。在谷歌PAML讨论站点上经常被问到的一个问题是是否存在并行版本。鉴于从笔记本电脑到超级计算机的多核架构无处不在,这种并行化将从根本上提高PAML分析的性能和可处理的数据规模。改进的计算性能将对PAML的所有现有用户非常有益,改进的用户界面可能有助于吸引新用户。该程序将根据GPL 3许可在其github网站(https://github.com/abacus-gene/paml)上发布。支持主要在谷歌讨论网站(https://groups.google.com/g/pamlsoftware)上提供,用户在那里发布和回答有关该软件的问题。PI定期访问该网站,特别是回答有关软件的更多技术用户问题。
英文摘要
PAML (for Phylogenetic Analysis by Maximum Likelihood) is a bioinformatics software tool widely used in the fields of molecular evolution, molecular phylogenetics, virology, biochemistry, and genomics. It is currently attracting over 1000 citations per year (with a total of >16K citations since its first release in 1993), and has a well-established user base. One of the major strengths of the package is its rich collection of sophisticated models for DNA and protein sequence evolution, which are useful in maximum likelihood and Bayesian methods in phylogenetics. The package can be used to compare different evolutionary trees, to infer adaptive molecular evolution affecting protein-coding genes, to reconstruct sequences in extinct ancestral species, etc. Despite being a widely used bioinformatics tool, PAML has a poor user interface and a steep learning curve. It is not parallelized, and is computationally inefficient when applied to large datasets. These issues have hindered its widespread adoption by new users and mean that existing users have to endure the computational burden. In this project, we propose to redesign and reimplement the key algorithms in the PAML programs, parallelize the code, and also develop a R interface. We will expand and improve its functionality by developing new mutation-selection models of codon evolution, which will lead to more accurate ancestral sequence reconstruction and more robust detection of genes under positive selection. Those improvements will greatly improve the usability and computational performance of the program package, making it an extremely valuable bioinformatics resource for the biosciences research community.This project is low-risk and high-impact. HPC technology, Pthreads and MPI have been around for many years. In phylogenetics, all the major likelihood and Bayesian programs have already been successfully parallelized, including PhyML, RAxML, MrBayes/RevBayes, and BEAST. Indeed PAML is uniquely bad in this sense for not having been parallelized. A frequently asked question at the google discussion site for PAML is whether a parallel version exists. Given the ubiquity of multicore architectures from laptops to supercomputers, this parallelisation will radically boost performance of PAML analysis and the scale of data that can be processed. The improved computational performance will be highly beneficial to all existing users of PAML and the improved user interface may help to attract new users. The program will be distributed at its github site (https://github.com/abacus-gene/paml) under the GPL 3 license. Support is mostly provided at the google discussion site (https://groups.google.com/g/pamlsoftware), where users post and answer questions about the software. The PI visits the site regularly, in particular, to answer more technical user queries about the software.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Efficient computational technologies to resolve the Timetree of Life: from ancient DNA to species-rich phylogenies
-
批准号:BB/Y004132/1
-
项目类别:Research Grant
-
资助金额:$64.4万
-
财政年份:2024
-
负责人:Ziheng Yang
-
依托单位:
NSFDEB-NERC: Integrating computational, phenotypic, and population-genomic approaches to reveal processes of cryptic speciation and gene flow in Madag
-
批准号:NE/X002071/1
-
项目类别:Research Grant
-
资助金额:$30.95万
-
财政年份:2023
-
负责人:Ziheng Yang
-
依托单位:
Bayesian inference of the mode of speciation and gene flow using genomic data
-
批准号:BB/X007553/1
-
项目类别:Research Grant
-
资助金额:$82.61万
-
财政年份:2023
-
负责人:Ziheng Yang
-
依托单位:
Bayesian implementation of the multispecies-coalescent-with-introgression (MSci) model for analysis of population genomic data
-
批准号:BB/T003502/1
-
项目类别:Research Grant
-
资助金额:$55.91万
-
财政年份:2020
-
负责人:Ziheng Yang
-
依托单位:
Efficient Bayesian phylogenomic dating with new models of trait evolution and rich diversities of living and fossil species
-
批准号:BB/T012951/1
-
项目类别:Research Grant
-
资助金额:$31.38万
-
财政年份:2020
-
负责人:Ziheng Yang
-
依托单位:
Phylogeographic inference using genomic sequence data under the multispecies coalescent model
-
批准号:BB/P006493/1
-
项目类别:Research Grant
-
资助金额:$50.8万
-
财政年份:2017
-
负责人:Ziheng Yang
-
依托单位:
Improving Bayesian methods for estimating divergence times integrating genomic and trait data
-
批准号:BB/N000609/1
-
项目类别:Research Grant
-
资助金额:$48.31万
-
财政年份:2016
-
负责人:Ziheng Yang
-
依托单位:
Statistical Methods for Genomic Analysis of Species Divergences
-
批准号:BB/K000896/1
-
项目类别:Research Grant
-
资助金额:$42.57万
-
财政年份:2013
-
负责人:Ziheng Yang
-
依托单位:
Bayesian Estimation of Species Divergence Times Integrating Fossil and Molecular Information
-
批准号:BB/J009709/1
-
项目类别:Research Grant
-
资助金额:$44.94万
-
财政年份:2012
-
负责人:Ziheng Yang
-
依托单位:
Representation and Incorporation of Fossil Data in Molecular Dating of Species Divergences
-
批准号:BB/G006431/1
-
项目类别:Research Grant
-
资助金额:$44.83万
-
财政年份:2009
-
负责人:Ziheng Yang
-
依托单位:
海外基金