Hierarchical phylogenetic models for analyzing multipartite sequence data

Hierarchical phylogenetic models for analyzing multipartite sequence data
复制标题

DOI:
10.1080/10635150390238879
复制
发表时间:
2003-10-01
期刊:
影响因子:
6.5
通讯作者:
Weiss, RE
Weiss, RE
中科院分区:
生物学1区
文献类型:
--
作者:
Suchard, MA;Kitchen, CMR;Weiss, RE

文献摘要

被引文献

相似文献

关于如何将多组分序列数据的信息纳入系统发育分析存在争议。严格的组合数据方法主张将所有分区连接起来,并估计一个进化历史,从而最大限度地提高数据的解释力。共识/独立方法支持两步程序,其中独立分析分区,然后从多个结果确定共识。严格的组合数据方法和先验独立参数的模型空间的混合是集成这些方法的流行方法。我们提出了一个替代的中间地带,通过构建贝叶斯层次系统发育模型。我们的分层框架使研究人员能够跨数据分区汇集信息,以提高单个分区的估计精度,同时允许估计和测试跨分区数量的趋势。这样的跨分区量包括从中绘制与分区内的序列相关的各个拓扑的分布。我们提出了标准的分层先验的连续进化参数跨分区,而拓扑结构的变化取决于研究问题。我们用三个例子来说明我们的模型。我们首先探索豚鼠(Cavia porcellus)的进化历史,使用13个线粒体基因的比对。分层模型返回比独立参数方法更精确的连续参数估计,而不会丢失数据的显着特征。第二,我们使用50个原核基因分析水平基因转移的频率。我们假设一个未知的物种级拓扑结构,并允许个别基因拓扑结构与此不同的一个小的可估计的概率。同时推断物种和个体基因拓扑结构返回17%的转移频率。我们还研究了从HIV+患者纵向采样的HIV序列。我们想知道治疗后CCR 5辅助受体病毒的发展是否代表了从疾病中期CXCR 4病毒的协同进化,或者最初感染CCR 5病毒的重新出现。分层模型通过假设每个患者的拓扑结构来自概率未知的多项分布,从多个不相关患者中合并分区。初步结果表明是进化而不是重新出现。
Debate exists over how to incorporate information from multipartite sequence data in phylogenetic analyses. Strict combined-data approaches argue for concatenation of all partitions and estimation of one evolutionary history, maximizing the explanatory power of the data. Consensus/independence approaches endorse a two-step procedure where partitions are analyzed independently and then a consensus is determined from the multiple results. Mixtures across the model space of a strict combined-data approach and a priori independent parameters are popular methods to integrate these methods. We propose an alternative middle ground by constructing a Bayesian hierarchical phylogenetic model. Our hierarchical framework enables researchers to pool information across data partitions to improve estimate precision in individual partitions while permitting estimation and testing of tendencies in across-partition quantities. Such across-partition quantities include the distribution from which individual topologies relating the sequences within a partition are drawn. We propose standard hierarchical priors on continuous evolutionary parameters across partitions, while the structure on topologies varies depending on the research problem. We illustrate our model with three examples. We first explore the evolutionary history of the guinea pig (Cavia porcellus) using alignments of 13 mitochondrial genes. The hierarchical model returns substantially more precise continuous parameter estimates than an independent parameter approach without losing the salient features of the data. Second, we analyze the frequency of horizontal gene transfer using 50 prokaryotic genes. We assume an unknown species-level topology and allow individual gene topologies to differ from this with a small estimable probability. Simultaneously inferring the species and individual gene topologies returns a transfer frequency of 17%. We also examine HIV sequences longitudinally sampled from HIV+ patients. We ask whether posttreatment development of CCR5 coreceptor virus represents concerted evolution from middisease CXCR4 virus or reemergence of initial infecting CCR5 virus. The hierarchical model pools partitions from multiple unrelated patients by assuming that the topology for each patient is drawn from a multinomial distribution with unknown probabilities. Preliminary results suggest evolution and not reemergence.