Sequencing facility and DNA source associated patterns of virus-mappable reads in whole-genome sequencing data.

Sequencing facility and DNA source associated patterns of virus-mappable reads in whole-genome sequencing data.
复制标题

在全基因组测序数据中,测序设施和DNA源相关的病毒式读取。

DOI:
10.1016/j.ygeno.2020.12.004
复制
发表时间:
2021-01
期刊:
影响因子:
4.4
通讯作者:
Li D
Li D
中科院分区:
生物学3区
文献类型:
--
作者:
Chen X;Li D

文献摘要

参考文献

被引文献

相似文献

在人类血液全基因组测序(WGS)数据中已经报道了许多病毒序列。然而,目前尚不清楚在多大程度上病毒可映射读取代表真正的病毒序列,而不是随机映射或来自样品制备、测序过程或其他来源的噪声。鉴定病毒可映射序列的模式可能为评估这些病毒序列的起源产生新的指标。我们鉴定了人类WGS数据集中成对端未映射的reads和与病毒参考文献一致的reads,然后比较了产生这些数据集的DNA来源和测序设施之间病毒可映射的reads的模式。然后,我们检查了与源和设施相关的病毒读取的潜在起源。在7个测序设备中,干净未映射reads的比例有显著差异(P < 2×10−16)。我们在2535个样本中鉴定了260,339个可映射到99个病毒参考文献的reads。大多数(86.7%)的病毒可映射reads(对应47个病毒参考文献),根据其不同的模式可分为四组,与测序设施或DNA来源密切相关(调整P值< 0.01)。这些reads的可能来源包括文库制备中的人工序列,细胞培养中的重组载体,以及与其宿主细菌共污染的噬菌体。在相同设施产生的其他数据集中反复观察到与测序设施相关的病毒可映射的读取和模式。我们构建了一个分析框架,并分析了可映射到病毒参考的未映射读段。该结果为深度测序数据中与测序设备和DNA源相关的批处理效应提供了新的认识,并可能有助于改进reads的生物信息学过滤。
Numerous viral sequences have been reported in the whole-genome sequencing (WGS) data of human blood. However, it is not clear to what degree the virus-mappable reads represent true viral sequences rather than random-mapping or noise originating from sample preparation, sequencing processes, or other sources. Identification of patterns of virus-mappable reads may generate novel indicators for evaluating the origins of these viral sequences. We characterized paired-end unmapped reads and reads aligned to viral references in human WGS datasets, then compared patterns of the virus-mappable reads among DNA sources and sequencing facilities which produced these datasets. We then examined potential origins of the source- and facility-associated viral reads. The proportions of clean unmapped reads among the seven sequencing facilities were significantly different (P < 2×10−16). We identified 260,339 reads that were mappable to a total of 99 viral references in 2,535 samples. The majority (86.7%) of these virus-mappable reads (corresponding to 47 viral references), which can be classified into four groups based on their distinct patterns, were strongly associated with sequencing facility or DNA source (adjusted P value < 0.01). Possible origins of these reads include artificial sequences in library preparation, recombinant vectors in cell culture, and phages co-contaminated with their host bacteria. The sequencing facility-associated virus-mappable reads and patterns were repeatedly observed in other datasets produced in the same facilities. We have constructed an analytic framework and profiled the unmapped reads mappable to viral references. The results provide a new understanding of sequencing facility- and DNA source-associated batch effects in deep sequencing data and may facilitate improved bioinformatics filtering of reads.
DOI: 10.1038/nprot.2007.135
发表时间: 2007-01-01
期刊: NATURE PROTOCOLS
影响因子: 14.8
作者:
Luo, Jinyong;Deng, Zhong-Liang;He, Tong-Chuan
通讯作者: He, Tong-Chuan
DOI: 10.1093/bioinformatics/bty595
发表时间: 2018-09-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Andrusch A;Dabrowski PW;Klenner J;Tausch SH;Kohl C;Osman AA;Renard BY;Nitsche A
通讯作者: Nitsche A
DOI: 10.1371/journal.pone.0097876
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Laurence M;Hatzis C;Brash DE
通讯作者: Brash DE
DOI: 10.1016/j.virol.2017.10.017
发表时间: 2018-01-01
期刊: Virology
影响因子: 3.7
作者:
Cantalupo PG;Katz JP;Pipas JM
通讯作者: Pipas JM
DOI: 10.1002/jcb.26717
发表时间: 2018-06-01
影响因子: 4
作者:
Cao, Jian;Li, Dawei
通讯作者: Li, Dawei