DNA barcode analysis: a comparison of phylogenetic and statistical classification methods.

DNA barcode analysis: a comparison of phylogenetic and statistical classification methods.
复制标题

DOI:
10.1186/1471-2105-10-s14-s10
复制
发表时间:
2009-11-10
期刊:
影响因子:
3
通讯作者:
Laredo C
Laredo C
中科院分区:
生物学4区
文献类型:
--
作者:
Austerlitz F;David O;Schaeffer B;Bleakley K;Olteanu M;Leblois R;Veuille M;Laredo C

文献摘要

被引文献

相似文献

DNA条形码的目的是根据个体在一个小位点的序列将其分配到给定的物种,通常是CO 1线粒体基因的一部分。除其他问题外,这提出了如何处理种内遗传变异和潜在的跨种多态性的问题。在这种情况下,我们研究了几种分配方法属于两个主要类别:(一)系统发育方法(邻居加入和PhyML),试图占DNA进化的谱系框架和(ii)监督分类方法(k-最近邻,CART,随机森林和内核方法)。这些方法从基本的到复杂的都有。我们研究了每种方法的能力,正确分类查询序列从相关物种的样本使用模拟和真实的数据。模拟数据集使用合并模拟生成,其中我们改变了系谱历史,突变参数,样本大小和物种数量。没有一种方法在所有情况下都是最好的。在所有方法中,最简单的方法“一个最近邻”被认为是最可靠的,因为数据集的参数发生了变化。最影响各种方法性能的参数是数据的分子多样性。添加遗传上独立的基因座-核基因-提高了大多数方法的预测性能。这项研究表明,分类学家可以通过选择最适合其样本结构的方法,或者在给定的方法下,增加样本量或改变分子多样性的数量来影响其分析的质量。这可以通过测序更多的mtDNA或通过测序额外的核基因来实现。在后一种情况下,他们可能还必须修改其数据分析方法。
DNA barcoding aims to assign individuals to given species according to their sequence at a small locus, generally part of the CO1 mitochondrial gene. Amongst other issues, this raises the question of how to deal with within-species genetic variability and potential transpecific polymorphism. In this context, we examine several assignation methods belonging to two main categories: (i) phylogenetic methods (neighbour-joining and PhyML) that attempt to account for the genealogical framework of DNA evolution and (ii) supervised classification methods (k-nearest neighbour, CART, random forest and kernel methods). These methods range from basic to elaborate. We investigated the ability of each method to correctly classify query sequences drawn from samples of related species using both simulated and real data. Simulated data sets were generated using coalescent simulations in which we varied the genealogical history, mutation parameter, sample size and number of species. No method was found to be the best in all cases. The simplest method of all, "one nearest neighbour", was found to be the most reliable with respect to changes in the parameters of the data sets. The parameter most influencing the performance of the various methods was molecular diversity of the data. Addition of genetically independent loci - nuclear genes - improved the predictive performance of most methods. The study implies that taxonomists can influence the quality of their analyses either by choosing a method best-adapted to the configuration of their sample, or, given a certain method, increasing the sample size or altering the amount of molecular diversity. This can be achieved either by sequencing more mtDNA or by sequencing additional nuclear genes. In the latter case, they may also have to modify their data analysis method.