Comparison of metatranscriptomic samples based on k-tuple frequencies.

Comparison of metatranscriptomic samples based on k-tuple frequencies.
复制标题

基于 k 元组频率的宏转录组样本比较

DOI:
10.1371/journal.pone.0084348
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Sun F
Sun F
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Wang Y;Liu L;Chen L;Chen T;Sun F

文献摘要

参考文献

被引文献

相似文献

背景样本的比较,或称β多样性,是生态学研究中的基本问题之一。下一代测序(NGS)技术使得在许多微生物群落中获得大量的元基因组和元转录短阅读序列成为可能。短片段的从头组装可能特别具有挑战性,因为基因组及其序列的数量通常是未知的,并且每个基因组的覆盖率可能非常低,其中不能使用传统的基于比对的序列比较方法。另一方面,基于k-字节组频率的无比对方法已经在元基因组样本的比较中产生了有希望的结果。然而,目前还不知道这些方法是否可以用于元翻译组数据集的比较,以及哪些不同的衡量标准执行得最好。结果我们将几个基于k-字节组频率的β多样性度量应用于来自PirroSequence454和Illumina测序平台的真实元翻译数据集,以评估它们对元翻译样本聚类的有效性,其中包括三个相异度度量、一个CVTree中的相异度度量、一个基于相对熵的度量S2和三个经典距离。结果表明,对于454和Illumina数据集,该方法能够在不同测序深度下将元基因组样本聚为不同的组,恢复影响微生物样本的环境梯度,对共存的元基因组数据集和元基因组数据集进行分类,并对测序错误具有较强的鲁棒性。我们还研究了元组大小和顺序对背景马尔可夫模型的影响。建立了用于实施所有分析步骤的软件管道,并可在http://code.google.com/p/d2-tools/.上获得结论基于k-字节组的序列签名方法能有效地揭示NGS阅读物中元翻译样本之间的主要群体和梯度差异。相异度度量在所有应用场景中都表现良好,并且其性能相对于马尔可夫模型的元组大小和阶数而言是稳健的。
Background The comparison of samples, or beta diversity, is one of the essential problems in ecological studies. Next generation sequencing (NGS) technologies make it possible to obtain large amounts of metagenomic and metatranscriptomic short read sequences across many microbial communities. De novo assembly of the short reads can be especially challenging because the number of genomes and their sequences are generally unknown and the coverage of each genome can be very low, where the traditional alignment-based sequence comparison methods cannot be used. Alignment-free approaches based on k-tuple frequencies, on the other hand, have yielded promising results for the comparison of metagenomic samples. However, it is not known if these approaches can be used for the comparison of metatranscriptome datasets and which dissimilarity measures perform the best. Results We applied several beta diversity measures based on k-tuple frequencies to real metatranscriptomic datasets from pyrosequencing 454 and Illumina sequencing platforms to evaluate their effectiveness for the clustering of metatranscriptomic samples, including three dissimilarity measures, one dissimilarity measure in CVTree, one relative entropy based measure S2 and three classical distances. Results showed that the measure can achieve superior performance on clustering metatranscriptomic samples into different groups under different sequencing depths for both 454 and Illumina datasets, recovering environmental gradients affecting microbial samples, classifying coexisting metagenomic and metatranscriptomic datasets, and being robust to sequencing errors. We also investigated the effects of tuple size and order of the background Markov model. A software pipeline to implement all the steps of analysis is built and is available at http://code.google.com/p/d2-tools/. Conclusions The k-tuple based sequence signature measures can effectively reveal major groups and gradient variation among metatranscriptomic samples from NGS reads. The dissimilarity measure performs well in all application scenarios and its performance is robust with respect to tuple size and order of the Markov model.
DOI: 10.1093/bioinformatics/btq365
发表时间: 2010-09-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Balzer S;Malde K;Lanzén A;Sharma A;Jonassen I
通讯作者: Jonassen I
DOI: 10.1016/j.gca.2009.07.039
发表时间: 2009-11-01
影响因子: 5
作者:
Dick, Gregory J.;Clement, Brian G.;Tebo, Bradley M.
通讯作者: Tebo, Bradley M.
DOI: 10.1371/journal.pone.0015545
发表时间: 2010-11-29
期刊: PloS one
影响因子: 3.7
作者:
Gilbert JA;Field D;Swift P;Thomas S;Cummings D;Temperton B;Weynberg K;Huse S;Hughes M;Joint I;Somerfield PJ;Mühling M
通讯作者: Mühling M
马尔可夫模型加上 k 字分布:产生用于序列比较的新颖统计测量的协同作用
DOI: 10.1093/bioinformatics/btn436
发表时间: 2008-10-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Dai, Qi;Yang, Yanchun;Wang, Tianming
通讯作者: Wang, Tianming
DOI: 10.1073/pnas.83.14.5155
发表时间: 1986-07-01
影响因子: 11.1
作者:
BLAISDELL, BE
通讯作者: BLAISDELL, BE