A Systematic Evaluation of High-Throughput Sequencing Approaches to Identify Low-Frequency Single Nucleotide Variants in Viral Populations.

A Systematic Evaluation of High-Throughput Sequencing Approaches to Identify Low-Frequency Single Nucleotide Variants in Viral Populations.
复制标题

DOI:
10.3390/v12101187
复制
发表时间:
2020-10-20
期刊:
Viruses
影响因子:
--
通讯作者:
Laing E
Laing E
中科院分区:
其他
文献类型:
--
作者:
King DJ;Freimanis G;Lasecka-Dykes L;Asfor A;Ribeca P;Waters R;King DP;Laing E

文献摘要

参考文献

被引文献

相似文献

高通量测序(例如 Illumina 提供的测序)是了解病毒群体内序列变异的有效方法。然而,在区分过程引入的错误和生物变异方面存在挑战,这极大地影响了我们识别次共识单核苷酸变异(SNV)的能力。在这里,我们采用系统方法来评估实验室和生物信息学管道,以准确识别病毒群体中的低频 SNV。人工 DNA 和 RNA“群体”是通过将已知的 SNV 以预定频率引入模板核酸中而创建的,然后在 Illumina MiSeq 平台上进行测序。这些用于评估丰度和起始输入材料类型、技术重复、读长和质量、短读对齐器以及百分比频率阈值对准确识别变体的能力的影响。分析显示,输入核酸的丰度和类型对 SNV 调用的准确性影响最大(通过微平均 Matthews 相关系数得分测量),DNA 和高 RNA 输入(107 个拷贝)允许以 0.2% 的频率调用变体。减少的 RNA 输入(105 个拷贝)需要更多的技术重复才能保持准确性,而低 RNA 输入(103 个拷贝)会出现共识级别的错误。还确定了在所有技术重复中识别的特定基序中识别的碱基错误,可以排除这些错误以进一步提高 SNV 调用准确性。这些发现表明,RNA 输入量低的样本应排除在 SNV 调用之外,并强调了优化用于准确识别序列变异的流程中的技术和生物信息学步骤的重要性。
High-throughput sequencing such as those provided by Illumina are an efficient way to understand sequence variation within viral populations. However, challenges exist in distinguishing process-introduced error from biological variance, which significantly impacts our ability to identify sub-consensus single-nucleotide variants (SNVs). Here we have taken a systematic approach to evaluate laboratory and bioinformatic pipelines to accurately identify low-frequency SNVs in viral populations. Artificial DNA and RNA “populations” were created by introducing known SNVs at predetermined frequencies into template nucleic acid before being sequenced on an Illumina MiSeq platform. These were used to assess the effects of abundance and starting input material type, technical replicates, read length and quality, short-read aligner, and percentage frequency thresholds on the ability to accurately call variants. Analyses revealed that the abundance and type of input nucleic acid had the greatest impact on the accuracy of SNV calling as measured by a micro-averaged Matthews correlation coefficient score, with DNA and high RNA inputs (107 copies) allowing for variants to be called at a 0.2% frequency. Reduced input RNA (105 copies) required more technical replicates to maintain accuracy, while low RNA inputs (103 copies) suffered from consensus-level errors. Base errors identified at specific motifs identified in all technical replicates were also identified which can be excluded to further increase SNV calling accuracy. These findings indicate that samples with low RNA inputs should be excluded for SNV calling and reinforce the importance of optimising the technical and bioinformatics steps in pipelines that are used to accurately identify sequence variants.
DOI: 10.1371/journal.pone.0176522
发表时间: 2017
期刊: PloS one
影响因子: 3.7
作者:
Operario DJ;Koeppel AF;Turner SD;Bao Y;Pholwat S;Banu S;Foongladda S;Mpagama S;Gratz J;Ogarkov O;Zhadova S;Heysell SK;Houpt ER
通讯作者: Houpt ER
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.3390/genes10080561
发表时间: 2019-08-01
期刊: GENES
影响因子: 3.5
作者:
Ferretti, Luca;Tennakoon, Chandana;Ribeca, Paolo
通讯作者: Ribeca, Paolo
DOI: 10.1099/0022-1317-80-8-1911
发表时间: 1999-08-01
影响因子: 3.8
作者:
Ellard, FM;Drew, J;King, AMQ
通讯作者: King, AMQ
DOI: 10.1371/journal.pone.0012303
发表时间: 2010-08-20
期刊: PloS one
影响因子: 3.7
作者:
Fischer W;Ganusov VV;Giorgi EE;Hraber PT;Keele BF;Leitner T;Han CS;Gleasner CD;Green L;Lo CC;Nag A;Wallstrom TC;Wang S;McMichael AJ;Haynes BF;Hahn BH;Perelson AS;Borrow P;Shaw GM;Bhattacharya T;Korber BT
通讯作者: Korber BT