Robust species taxonomy assignment algorithm for 16S rRNA NGS reads: application to oral carcinoma samples.

Robust species taxonomy assignment algorithm for 16S rRNA NGS reads: application to oral carcinoma samples.
复制标题

DOI:
10.3402/jom.v7.28934
复制
发表时间:
2015
影响因子:
4.5
通讯作者:
Chen T
Chen T
中科院分区:
医学2区
文献类型:
--
作者:
Al-Hebshi NN;Nasher AT;Idris AM;Chen T

文献摘要

被引文献

相似文献

由于无法将reads分类到物种水平,下一代测序(NGS)在评估与口腔鳞状细胞癌(OSCC)相关的细菌方面的有效性受到了破坏。本研究的目的是开发一种强大的算法,用于口腔样本中NGS reads的物种水平分类,并对其进行中试测试,以分析OSCC组织中的细菌。从三个OSCC DNA样本中制备细菌16S V1-V3文库,并使用454的FLX化学方法进行测序。使用一种新颖的多阶段算法对≥350bp的高质量,良好对齐和非嵌合的reads进行分类,该算法包括将reads与修订版Human Oral Microbiome Database (HOMD), HOMDEXT (HOMDEXT)和Greengene Gold (GGG)中的参考序列进行比对,比对覆盖率和百分比一致性≥98%,然后根据最高命中参考序列分配到物种水平。首先是在HOMD中命中,然后是HOMDEXT,最后是GGG。不匹配的reads需要进行操作分类单元分析。近92.8%的reads与updated-HOMD 13.2匹配,1.83%与trusted-HOMDEXT匹配,1.36%与modified-GGG匹配。在所有匹配的reads中,99.6%被分类到物种水平。共鉴定出11门228个种级分类群;最丰富的是变形菌门、拟杆菌门、厚壁菌门、梭菌门和放线菌门。所有样本共检测到35个种级分类群。平均而言,口腔普氏菌、黄奈瑟菌、黄奈瑟菌/亚黄奈瑟菌、多形核梭菌、segnis聚集菌、mitis链球菌和牙周梭菌数量最多。在两个样本中检测到脆弱拟杆菌,这是一种很少从口腔中分离出来的物种。该多阶段算法最大限度地提高了分类到物种水平的reads的比例,同时通过优先考虑人类,口头参考集来确保可靠的分类。将该算法应用于OSCC样本显示出较高的多样性。除了口腔分类群外,还发现了一些人类非口腔分类群,其中一些在口腔中很少发现。
Usefulness of next-generation sequencing (NGS) in assessing bacteria associated with oral squamous cell carcinoma (OSCC) has been undermined by inability to classify reads to the species level. The purpose of this study was to develop a robust algorithm for species-level classification of NGS reads from oral samples and to pilot test it for profiling bacteria within OSCC tissues. Bacterial 16S V1-V3 libraries were prepared from three OSCC DNA samples and sequenced using 454's FLX chemistry. High-quality, well-aligned, and non-chimeric reads ≥350 bp were classified using a novel, multi-stage algorithm that involves matching reads to reference sequences in revised versions of the Human Oral Microbiome Database (HOMD), HOMD extended (HOMDEXT), and Greengene Gold (GGG) at alignment coverage and percentage identity ≥98%, followed by assignment to species level based on top hit reference sequences. Priority was given to hits in HOMD, then HOMDEXT and finally GGG. Unmatched reads were subject to operational taxonomic unit analysis. Nearly, 92.8% of the reads were matched to updated-HOMD 13.2, 1.83% to trusted-HOMDEXT, and 1.36% to modified-GGG. Of all matched reads, 99.6% were classified to species level. A total of 228 species-level taxa were identified, representing 11 phyla; the most abundant were Proteobacteria, Bacteroidetes, Firmicutes, Fusobacteria, and Actinobacteria. Thirty-five species-level taxa were detected in all samples. On average, Prevotella oris, Neisseria flava, Neisseria flavescens/subflava, Fusobacterium nucleatum ss polymorphum, Aggregatibacter segnis, Streptococcus mitis, and Fusobacterium periodontium were the most abundant. Bacteroides fragilis, a species rarely isolated from the oral cavity, was detected in two samples. This multi-stage algorithm maximizes the fraction of reads classified to the species level while ensuring reliable classification by giving priority to the human, oral reference set. Applying the algorithm to OSCC samples revealed high diversity. In addition to oral taxa, a number of human, non-oral taxa were also identified, some of which are rarely detected in the oral cavity.