Fast supertree construction using quartet joining
Fast supertree construction using quartet joining
批准号:
BB/G024707/1
负责人:
Peter Foster
金额:
$15.5万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2009
资助国家:
英国
项目状态:
已结题
起止时间:
2009 至 --
中文摘要
所有曾经生活过的生物都是通过共同的祖先和后代在一棵生命树上联系在一起的,这是生物科学的主要见解之一。对这些系统发育关系的了解有助于科学家理解我们今天看到的生命的巨大多样性是如何起源的,为推断生物是如何进化的提供了一个框架,并允许测试试图解释这种多样性并确定产生这种多样性的机制的假说。系统发育关系可以用形态学来推断,但越来越多地从DNA或氨基酸序列数据来推断。然而,由于推断中的错误或因为基因树与物种树不同,推测的单个基因的系统发育可能与真实的物种系统发育不同(不一致)。例如,当基因在物种之间水平转移时,例如在一些细菌产生抗生素耐药性的过程中就发生了这种情况,或者当基因复制并随后丢失时,可能会出现后一种情况。这就提出了如何最好地进行系统基因组学(基因组规模数据的系统发生分析)的问题,目前正在寻求两种替代策略(1)将所有基因组合成单一分析和(2)构建超树--单个基因树的合成。超树方法可以被认为是一种“分而治之”的方法,在这种方法中,一个大的系统发育问题被分解成更小的问题,然后这些问题被组合在一起,给出一个全局解决方案。支持这一点的是这样一种期望,即单个基因树可以更容易或更有效地进行分析,因为它们更小,而且它们只包括那些特定基因可用的分类群。这还假设可以有效地组合各个树中的信息,但不幸的是,当前在实践中最依赖的超树方法具有许多明显不受欢迎的属性,例如产生与每个输入树的真关系相矛盾的超树(因此,如果任何输入树为真,则该关系必须为真)。我们建议开发一种新的超树方法,该方法使用逻辑推理从基因树集合中构建物种系统发育图,并在软件中实现,并用模拟和经验数据进行测试。在这种方法中,通过添加叶子来生长超树;在输入的树中,关于在哪里放置新叶子的推论是由可以被认为是系统发育信息量子的四元组给出的。需要新的方法使研究人员能够最大限度地利用快速增长的完整基因组序列,这些基因组序列可能与了解代谢途径的演变、耐药性、药物发现、流行病学和与历史气候变化有关的多样化研究有关。随着技术的进步,新的基因组数据的生产速度大大提高;现在一个下午就可以生产出原核生物的完整基因组。现在需要改进用于分析这一海量数据的方法,目标是用更有根据的替代方法取代特别方法。基于其逻辑基础、灵活性和计算速度,我们预计这将是系统发育分析的一种选择方法,但这需要通过模拟来证实,以显示其性质和确定其错误率,并通过经验测试提供概念证明。
英文摘要
That all kinds of organisms that have ever lived are related through common ancestry and descent in one Tree of Life is one of the major insights of bological science. Knowledge of these phylogenetic relationships helps scientists to understand how the great diversity of life we see today has originated, provides a framework for inferring how living things have evolved, and allows testing hypotheses that seek to explain this diversity and identify the mechanisms that have generated it. Phylogenetic relationships can be inferred using morphology but are increasingly inferred from DNA or amino acid sequence data. However, the inferred phylogeny of a single gene may differ from (be incongruent with) the true species phylogeny, either due to errors in the inference or because the gene tree is not identical to the species tree. The latter can arise when, for example, genes are transferred horizontally between species, as has happened in the development of antibiotic resistance in some bacteria, or when genes are duplicated and subsequently lost. This raises questions of how best to do phylogenomics (the phylogenetic analysis of genomic scale data) with two alternative strategies currently being pursued (1) combining all genes into a single analysis and (2) building a supertree - a synthesis of the individual gene trees. Supertree methods can be considered a 'divide-and-conquer' approach where a large phylogenetic problem is decomposed into smaller problems which are then combined to give a global solution. Underpinning this is the expectation that individual gene trees can be more easily or effectively analysed because they are smaller and because they include only those taxa for which particular genes are available. This also assumes that the information in the individual trees can be combined efficiently, but unfortunately the supertree methods that are currently most relied upon in practice have a number of obviously undesirable properties, such as producing supertrees that contradict relationships that are true of every input tree (and which therefore must be true if any input tree is true). We propose to develop a new supertree method that uses logical inference to make species phylogenies from collections of gene trees, to implement it in software, and to test it with simulations and empirical data. In this method a supertree is grown by adding leaves; the inference about where to put new leaves is given by 'quartets', which can be considered the quanta of phylogenetic information, in the input trees. The new method is needed to enable researchers to make best use of the rapidly expanding number of complete genome sequences which may be of relevance to understanding the evolution of metabolic pathways, of drug resistance, to drug discovery, epidemiology, and diversification studies linked to historical climate change. Technical advances have seen the massive increases in the rate of production of new genomic data; complete genomes of prokaryotes can now be produced in an afternoon. Advances are now needed in the methods used to analyse this flood of data, and aim to replace ad hoc methods with better-founded alternatives. Based on its logical foundation, its flexibility, and on the speed of its computation, we expect that this will be a method of choice in phylogenomic analysis, but this needs to be confirmed through simulation to show its properties and determine its error rates, and through empirical tests that will provide proof of concept.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1073/pnas.1618463114
发表时间:
2017-06-06
期刊:
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA
影响因子:
11.1
作者:
[Williams, Tom A., Szollosi, Gergely J., Embley, T. Martin]
通讯作者:
Embley, T. Martin
DOI:
10.1098/rsos.140436
发表时间:
2015-08
期刊:
Royal Society open science
影响因子:
3.5
作者:
[Akanni WA, Wilkinson M, Creevey CJ, Foster PG, Pisani D]
通讯作者:
Pisani D
DOI:
10.1186/1471-2105-15-183
发表时间:
2014-06-12
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Akanni WA, Creevey CJ, Wilkinson M, Pisani D]
通讯作者:
Pisani D
How do eukaryotic CO2 fixers co-exist with faster growing prokaryotic CO2 fixers in the oligotrophic ocean covering 40% of Earth?
-
批准号:NE/M015831/1
-
项目类别:Research Grant
-
资助金额:$17.22万
-
财政年份:2015
-
负责人:Peter Foster
-
依托单位:
海外基金