Confounding factors in profiling of locus-specific human endogenous retrovirus (HERV) transcript signatures in primary T cells using multi-study-derived datasets.

Confounding factors in profiling of locus-specific human endogenous retrovirus (HERV) transcript signatures in primary T cells using multi-study-derived datasets.
复制标题

DOI:
10.1186/s12920-023-01486-y
复制
发表时间:
2023-04-03
影响因子:
2.7
通讯作者:
--
中科院分区:
医学3区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

人类内源性逆转录病毒(HERV)是重复序列元素,是人类基因组的重要组成部分。它们在发育中的作用已被充分记录,现在有越来越多的证据表明,HERV表达失调也会导致各种人类疾病。虽然对HERV基因的研究过去一直受到其高序列相似性的阻碍,但先进的测序技术和分析工具使该领域得到了发展。我们现在第一次能够进行位点特异性HERV分析,破译这些元件的表达模式、调控网络和生物学功能。为了做到这一点,我们不可避免地依赖于通过公共领域提供的组学数据集。然而,技术参数不可避免地存在差异,使得研究间分析具有挑战性。我们在这里使用来自多个来源的数据集来解决分析位点特异性HERV转录组的混杂因素问题。我们收集了CD4和CD8原代T细胞的RNAseq数据集,并提取了3220个元素的HERV表达谱,类似于大多数完整的、接近全长的原病毒。考虑到测序参数和批处理效应,我们比较了不同数据集的HERV特征,并确定了多源数据中HERV表达分析的允许特征。我们可以证明,考虑到测序参数,测序深度对HERV特征结果影响最大。测序样品更深入地拓宽了表达HERV元件的谱。测序模式和读取长度是次要参数。然而,我们发现来自较小的RNAseq数据集的HERV特征确实可靠地揭示了最丰富表达的HERV元素。总的来说,样本和研究之间的HERV特征有很大的重叠,表明在CD4和CD8 T细胞中有强大的HERV转录物特征。此外,我们发现减少批效应的措施对于揭示细胞类型之间的基因和HERV表达差异至关重要。在这样做之后,HERV转录组在本体密切相关的CD4和CD8 T细胞之间的差异变得明显。在我们确定检测基因座特异性HERV表达的测序和分析参数的系统方法中,我们提供的证据表明,对来自多个研究的RNAseq数据集的分析有助于对生物学发现的信心。当生成新的HERV表达数据集时,我们建议与标准基因转录组管道相比增加序列深度(> = 100万个reads)。最后,需要实施批次效应减少措施,以便进行差异表达分析。在线版本包含补充材料,可在10.1186/s12920-023-01486-y获得。
Human endogenous retroviruses (HERV) are repetitive sequence elements and a substantial part of the human genome. Their role in development has been well documented and there is now mounting evidence that dysregulated HERV expression also contributes to various human diseases. While research on HERV elements has in the past been hampered by their high sequence similarity, advanced sequencing technology and analytical tools have empowered the field. For the first time, we are now able to undertake locus-specific HERV analysis, deciphering expression patterns, regulatory networks and biological functions of these elements. To do so, we inevitable rely on omics datasets available through the public domain. However, technical parameters inevitably differ, making inter-study analysis challenging. We here address the issue of confounding factors for profiling locus-specific HERV transcriptomes using datasets from multiple sources. We collected RNAseq datasets of CD4 and CD8 primary T cells and extracted HERV expression profiles for 3220 elements, resembling most intact, near full-length proviruses. Looking at sequencing parameters and batch effects, we compared HERV signatures across datasets and determined permissive features for HERV expression analysis from multiple-source data. We could demonstrate that considering sequencing parameters, sequencing-depth is most influential on HERV signature outcome. Sequencing samples deeper broadens the spectrum of expressed HERV elements. Sequencing mode and read length are secondary parameters. Nevertheless, we find that HERV signatures from smaller RNAseq datasets do reliably reveal most abundantly expressed HERV elements. Overall, HERV signatures between samples and studies overlap substantially, indicating a robust HERV transcript signature in CD4 and CD8 T cells. Moreover, we find that measures of batch effect reduction are critical to uncover genic and HERV expression differences between cell types. After doing so, differences in the HERV transcriptome between ontologically closely related CD4 and CD8 T cells became apparent. In our systematic approach to determine sequencing and analysis parameters for detection of locus-specific HERV expression, we provide evidence that analysis of RNAseq datasets from multiple studies can aid confidence of biological findings. When generating de novo HERV expression datasets we recommend increased sequence depth ( > = 100 mio reads) compared to standard genic transcriptome pipelines. Finally, batch effect reduction measures need to be implemented to allow for differential expression analysis. The online version contains supplementary material available at 10.1186/s12920-023-01486-y.
DOI: 10.1186/s13100-016-0080-x
发表时间: 2016
期刊: Mobile DNA
影响因子: 4.9
作者:
Babaian A;Mager DL
通讯作者: Mager DL
DOI: 10.1093/bioinformatics/btu638
发表时间: 2015-01-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Anders S;Pyl PT;Huber W
通讯作者: Huber W
DOI: 10.1186/s12920-015-0146-5
发表时间: 2015-11-03
影响因子: 2.7
作者:
Haase K;Mösch A;Frishman D
通讯作者: Frishman D
DOI: 10.1126/science.abk3112
发表时间: 2022-04
期刊: SCIENCE
影响因子: 56.9
作者:
Hoyt, Savannah J.;Storer, Jessica M.;Hartley, Gabrielle A.;Grady, Patrick G. S.;Gershman, Ariel;de Lima, Leonardo G.;Limouse, Charles;Halabian, Reza;Wojenski, Luke;Rodriguez, Matias;Altemose, Nicolas;Rhie, Arang;Core, Leighton J.;Gerton, Jennifer L.;Makalowski, Wojciech;Olson, Daniel;Rosen, Jeb;Smit, Arian F. A.;Straight, Aaron F.;Vollger, Mitchell R.;Wheeler, Travis J.;Schatz, Michael C.;Eichler, Evan E.;Phillippy, Adam M.;Timp, Winston;Miga, Karen H.;O'Neill, Rachel J.
通讯作者: O'Neill, Rachel J.
DOI: 10.3389/fchem.2017.00035
发表时间: 2017
影响因子: 5.5
作者:
Buzdin AA;Prassolov V;Garazha AV
通讯作者: Garazha AV