Minimum entropy decomposition: unsupervised oligotyping for sensitive partitioning of high-throughput marker gene sequences.

Minimum entropy decomposition: unsupervised oligotyping for sensitive partitioning of high-throughput marker gene sequences.
复制标题

DOI:
10.1038/ismej.2014.195
复制
发表时间:
2015-03-17
期刊:
The ISME journal
影响因子:
--
通讯作者:
Sogin ML
Sogin ML
中科院分区:
其他
文献类型:
--
作者:
Eren AM;Morrison HG;Lescault PJ;Reveillaud J;Vineis JH;Sogin ML

文献摘要

被引文献

相似文献

分子微生物生态学研究通常采用大的标记基因数据集,例如核糖体RNA,以代表微生物群落中单细胞基因组的出现。大规模并行DNA测序技术使得能够对标记基因文库进行广泛调查,这些标记基因文库有时包括几乎相同的序列。计算方法,依赖于成对序列比对的相似性评估和从头聚类与事实上的相似性阈值分区高通量测序数据集约束微生物群落的精细尺度分辨率的描述。最小熵分解(MED)提供了一种计算效率高的方法来划分标记基因数据集到“MED节点”,代表同质操作分类单元。通过采用香农熵,MED仅使用跨读段的信息丰富的核苷酸位置,并迭代地划分大数据集,同时省略随机变化。当应用于两种深海隐蔽海绵Hexadella dedritifera和Hexadella cf. dedritifera,MED解决了一个关键的Gammaproteobacteria集群到多个MED节点,是特定于不同的海绵,并揭示了这些密切相关的同域海绵物种保持不同的微生物群落。对之前发表的人类口腔微生物组数据集的MED分析还显示,序列变异小于1%的分类群分布在口腔中的不同生态位。MED算法背后的信息理论指导的分解过程使得能够灵敏地区分标记基因扩增子数据集中的密切相关的生物体,而不依赖于广泛的计算分析和用户监督。
Molecular microbial ecology investigations often employ large marker gene datasets, for example, ribosomal RNAs, to represent the occurrence of single-cell genomes in microbial communities. Massively parallel DNA sequencing technologies enable extensive surveys of marker gene libraries that sometimes include nearly identical sequences. Computational approaches that rely on pairwise sequence alignments for similarity assessment and de novo clustering with de facto similarity thresholds to partition high-throughput sequencing datasets constrain fine-scale resolution descriptions of microbial communities. Minimum Entropy Decomposition (MED) provides a computationally efficient means to partition marker gene datasets into ‘MED nodes', which represent homogeneous operational taxonomic units. By employing Shannon entropy, MED uses only the information-rich nucleotide positions across reads and iteratively partitions large datasets while omitting stochastic variation. When applied to analyses of microbiomes from two deep-sea cryptic sponges Hexadella dedritifera and Hexadella cf. dedritifera, MED resolved a key Gammaproteobacteria cluster into multiple MED nodes that are specific to different sponges, and revealed that these closely related sympatric sponge species maintain distinct microbial communities. MED analysis of a previously published human oral microbiome dataset also revealed that taxa separated by less than 1% sequence variation distributed to distinct niches in the oral cavity. The information theory-guided decomposition process behind the MED algorithm enables sensitive discrimination of closely related organisms in marker gene amplicon datasets without relying on extensive computational heuristics and user supervision.