A DNA Barcoding system integrating multigene sequence data

A DNA Barcoding system integrating multigene sequence data
复制标题

整合多基因序列数据的DNA条形码系统

DOI:
10.1111/2041-210x.12366
复制
发表时间:
2015-08-01
影响因子:
6.6
通讯作者:
Zhu, Chao-Dong
Zhu, Chao-Dong
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Chesters, Douglas;Zheng, Wei-Min;Zhu, Chao-Dong

文献摘要

被引文献

相似文献

已经开发了许多系统用于DNA序列数据的分类鉴定。然而,在真核生物中,这些系统主要基于单个预定义的基因,因此容易受到来自有限特征采样的偏差的影响,并且不能识别基因组起源的大多数序列。我们在这里演示了多基因DNA条形码的实现。首先,一个参考框架是建立频繁测序的基因座。然后通过切除与参考同源的序列并指定物种名称来组织查询序列数据,其中查询和参考之间的序列相似性水平福尔斯在通常观察到的种内变异的(基因适当的)水平内。该方法相比,一些现有的方法,包括'风笛_',一个重新实现的分类分配的植物。78%的物种和94%的已知存在于节肢动物测试查询的属被正确推断出所提出的多基因系统。最重要的是,物种识别率比仅使用COI方法有所提高。查询中24%的物种仅在非COI基因中发现,在许多其他基因座上物种分配的准确性没有明显降低。同样,使用非COI列对合并宏基因组数据集进行了额外的物种分配。在273个蜜蜂序列的较小查询数据集上,使用修改的距离计算进行物种分配的准确性与基于系统发育的分类识别没有区别。标准化的单片段DNA条形码仍然是PCR生成的序列数据物种鉴定的宝贵工具。本文开发的方法用其他基因组数据补充了已建立的物种密集DNA条形码骨架,通过整合独立的遗传基因座减少了错误,并允许对非条形码片段进行额外的鉴定。后者将特别适用于使用下一代测序平台监测社区基因组学。
A number of systems have been developed for taxonomic identification of DNA sequence data. However, in eukaryotes, these systems are largely based on single predefined genes, and thus are vulnerable to biases from limited character sampling, and are not able to identify most sequences of genomic origin. We here demonstrate an implementation for multigene DNA barcoding. First, a reference framework is built of frequently sequenced loci. Query sequence data are then organized by excising sequences homologous to references and assigning species names where the level of sequence similarity between query and reference falls within the (gene‐appropriate) level of intraspecific variation usually observed. The approach is compared to some existing methods including ‘bagpipe_phylo’, a re‐implementation for taxonomic assignment on phylogenies. Seventy‐eight per cent of the species and 94% of the genera known to be present in arthropod test queries were correctly inferred by the proposed multigene system. Most critically, the rate of species identification was increased over using a COI‐only approach. Twenty‐four per cent of species in the queries were found only in non‐COI genes, with no clear reduction in the accuracy of species assignment at many of these other loci. Similarly, additional species assignments were made for a pooled metagenomic data set using non‐COI columns. On a smaller query data set of 273 bee sequences, the accuracy of species assignment using modified calculation of distances was indistinguishable from phylogeny‐based taxonomic identification. Standardized single fragment DNA barcoding remains an invaluable tool in species identification for PCR‐generated sequence data. The approach developed here supplements the established species‐dense DNA barcode backbone with other genomic data, reducing error via integration of independent genetic loci and permitting additional identifications for non‐barcode fragments. The latter will be particularly relevant in monitoring of community genomics using next‐generation sequencing platforms.