Evaluation of High-Throughput Sequencing for Identifying Known and Unknown Viruses in Biological Samples

Evaluation of High-Throughput Sequencing for Identifying Known and Unknown Viruses in Biological Samples
复制标题

DOI:
10.1128/jcm.00850-11
复制
发表时间:
2011-09-01
影响因子:
9.4
通讯作者:
Eloit, Marc
Eloit, Marc
中科院分区:
医学2区
文献类型:
--
作者:
Cheval, Justine;Sauvage, Virginie;Eloit, Marc

文献摘要

被引文献

相似文献

高通量测序从未克隆的DNA中获得大量的短序列读段,并已迅速成为鉴定生物样品中病毒的主要工具,特别是当靶序列不确定时。在这项研究中,我们评估了基于Roche-454基因组测序仪或Illumina基因组分析仪平台检测生物样品中病毒的管道的分析灵敏度。我们对人工掺入各种病毒的生物样品进行了测序,这些病毒的基因组由单链或双链DNA或RNA组成,包括线性或环状单链DNA。以非常低的浓度加入病毒,通常相当于定量逆转录酶PCR(RT-PCR)检测验证水平的3倍或0.8倍。对于在公共核苷酸序列数据库中所代表的病毒或类似于所代表的病毒,我们表明Illumina的较高输出与更高的灵敏度相关,接近优化的定量(RT-)PCR。在这项盲法研究中,实现了病毒鉴定,未出现错误鉴定。然而,在这些低浓度下,由Illumina平台生成的读段的数量太小,以至于不能在不使用参考序列的情况下促进重叠群的组装,从而排除了未知病毒的检测。当病毒载量足够高时,从头组装允许产生对应于几乎全长基因组的长重叠群,因此应有助于鉴定新病毒。
High-throughput sequencing furnishes a large number of short sequence reads from uncloned DNA and has rapidly become a major tool for identifying viruses in biological samples, and in particular when the target sequence is undefined. In this study, we assessed the analytical sensitivity of a pipeline for detection of viruses in biological samples based on either the Roche-454 genome sequencer or Illumina genome analyzer platforms. We sequenced biological samples artificially spiked with a wide range of viruses with genomes composed of single or double-stranded DNA or RNA, including linear or circular single-stranded DNA. Viruses were added at a very low concentration most often corresponding to 3 or 0.8 times the validated level of detection of quantitative reverse transcriptase PCRs (RT-PCRs). For the viruses represented, or resembling those represented, in public nucleotide sequence databases, we show that the higher output of Illumina is associated with a much greater sensitivity, approaching that of optimized quantitative (RT-)PCRs. In this blind study, identification of viruses was achieved without incorrect identification. Nevertheless, at these low concentrations, the number of reads generated by the Illumina platform was too small to facilitate assembly of contigs without the use of a reference sequence, thus precluding detection of unknown viruses. When the virus load was sufficiently high, de novo assembly permitted the generation of long contigs corresponding to nearly full-length genomes and thus should facilitate the identification of novel viruses.