How does normalization impact RNA-seq disease diagnosis?

How does normalization impact RNA-seq disease diagnosis?
复制标题

标准化如何影响 RNA-seq 疾病诊断?

DOI:
10.1016/j.jbi.2018.07.016
复制
发表时间:
2018-09-01
影响因子:
4.5
通讯作者:
Men, Ke
Men, Ke
中科院分区:
医学3区
文献类型:
--
作者:
Han, Henry;Men, Ke

文献摘要

被引文献

相似文献

随着下一代高通量技术的兴起,RNA-seq数据在疾病诊断中发挥着越来越重要的作用,其中归一化被认为是产生可比较样本的必要程序。最近的研究已经看到不同的标准化方法提出,以消除各种技术偏差在RNA测序。然而,目前尚无研究评估正常化对RNA-seq疾病诊断的影响。在本研究中,我们通过分析结构化大数据来研究这个问题:从TCGA门户网站获得的RNA-seq数据,因为它在RNA-seq疾病诊断中很受欢迎。我们提出了一种新的归一化效果测试算法,诊断指数(d-index)和数据熵,通过使用最先进的机器学习模型来分析和评估归一化对RNA-seq疾病诊断的影响。此外,我们提出了一个原始的可视化分析来比较规范化数据与原始数据的性能。我们发现,规范化数据通常产生的诊断水平与原始数据相当,甚至更低。此外,一些标准化方法(如RPKM)甚至对疾病诊断产生负面影响。另一方面,原始数据似乎有可能更好地破译病理状态,或者至少比数据规范化时具有可比性。我们的可视化分析还表明,一些归一化方法甚至会带来“异常值”,这不可避免地降低了诊断中的样本可检测性。更重要的是,我们的数据熵分析表明,规范化数据通常显示出与原始数据相当或更低的熵值。具有高熵值的数据往往比具有低熵值的数据具有更好的诊断效果。此外,我们发现,在诊断过程中,高维不平衡(HDI)数据不受任何归一化过程的影响,并且几乎所有的机器学习模型都无法识别大多数类型,尽管是原始数据或归一化数据。我们的研究结果表明,规范化数据在疾病诊断方面可能不会比原始数据表现出统计学上的显著优势。这进一步表明,在RNA-seq疾病诊断中,正常化可能不是必不可少的程序,或者至少一些正常化过程可能不是。相反,在不同的病理条件下,原始数据可以更好地捕获更多的原始转录组模式。
With the surge of next generation high-throughput technologies, RNA-seq data is playing an increasingly important role in disease diagnosis, in which normalization is assumed as an essential procedure to produce comparable samples. Recent studies have seen different normalization methods proposed to remove various technical biases in RNA sequencing. However, there are no previous studies evaluating the impacts of normalization on RNA-seq disease diagnosis.In this study, we investigate this problem by analyzing structured big data: RNA-seq data acquired from the TCGA portal for its popularity in RNA-seq disease diagnosis. We propose a novel normalization effect test algorithm, diagnostic index (d-index), and data entropy to analyze and evaluate the impacts of normalization on RNA-seq disease diagnosis by using state-of-the-art machine learning models. Furthermore, we present an original visualization analysis to compare the performance of normalized data versus raw data.We have found that normalized data yields generally an equivalent or even lower level diagnosis than its raw data. Moreover, some normalization approaches (e.g. RPKM) even bring negative effects in disease diagnosis. On the other hand, raw data seems to have the potential to decipher pathological status better or at least comparable than when the data is normalized. Our visualization analysis also shows that some normalization methods even bring 'outliers', which unavoidably decreases sample detectability in diagnosis. More importantly, our data entropy analysis shows that normalized data usually demonstrates equivalent or lower entropy values than raw data. Those data with high entropy values tend to achieve better diagnosis than those with low entropy values. In addition, we found that high-dimensional imbalance (HDI) data is unaffected by any normalization procedures in diagnosis, and fails almost all machine learning models by only recognizing majority types in spite of raw or normalized data.Our results suggest that normalized data may not demonstrate statistically significant advantages in disease diagnosis than its raw form. It further implies that normalization may not be an indispensable procedure in RNA-seq disease diagnosis or at least some normalization processes may not be. Instead, raw data may perform better for capturing more original transcriptome patterns in different pathological conditions.