Determining the optimal number of independent components for reproducible transcriptomic data analysis.

Determining the optimal number of independent components for reproducible transcriptomic data analysis.
复制标题

DOI:
10.1186/s12864-017-4112-9
复制
发表时间:
2017-09-11
期刊:
影响因子:
4.4
通讯作者:
Zinovyev A
Zinovyev A
中科院分区:
生物学2区
文献类型:
--
作者:
Kairov U;Cantini L;Greco A;Molkenov A;Czerwinska U;Barillot E;Zinovyev A

文献摘要

参考文献

被引文献

相似文献

独立成分分析(ICA)是一种将基因表达数据建模为一组统计上独立的隐藏因素的作用的方法。ICA的输出取决于一个基本参数:要计算的组件(因子)的数量。在盲源分离技术应用于转录组学数据中,该参数的最佳选择与确定有效数据维数有关,仍然是一个悬而未决的问题。在这里,我们解决了优化转录组学数据分析中统计独立成分数量的问题,以便在多个ICA运行(在相同或不同有效维度内)和多个独立数据集中成分的可重复性。为此,我们根据独立成分在多次ICA计算中的稳定性对其进行排序,并定义了与稳定性曲线质变点相对应的不同数量的成分(最稳定转录组维,MSTD)。基于大量的数据,我们证明了ICA分解的生物学可解释性需要足够数量的维度,并且与不太稳定的成分相比,排名低于MSTD的最稳定成分在独立研究中有更多的机会被复制。与此同时,我们表明转录组学数据集可以减少到相对较高的维度,而不会失去ICA的可解释性,即使更高的维度会产生由小基因集驱动的组件。我们建议将ICA应用于转录组学数据,并根据其可重复性对组件进行优先排序,从而加强生物学解释。计算太少的组件(比MSTD少得多)对于结果的可解释性来说不是最佳的。在MSTD范围内排名的成分在独立研究中有更多的机会被复制。本文的在线版本(10.1186/s12864-017-4112-9)包含补充材料,授权用户可以使用。
Independent Component Analysis (ICA) is a method that models gene expression data as an action of a set of statistically independent hidden factors. The output of ICA depends on a fundamental parameter: the number of components (factors) to compute. The optimal choice of this parameter, related to determining the effective data dimension, remains an open question in the application of blind source separation techniques to transcriptomic data. Here we address the question of optimizing the number of statistically independent components in the analysis of transcriptomic data for reproducibility of the components in multiple runs of ICA (within the same or within varying effective dimensions) and in multiple independent datasets. To this end, we introduce ranking of independent components based on their stability in multiple ICA computation runs and define a distinguished number of components (Most Stable Transcriptome Dimension, MSTD) corresponding to the point of the qualitative change of the stability profile. Based on a large body of data, we demonstrate that a sufficient number of dimensions is required for biological interpretability of the ICA decomposition and that the most stable components with ranks below MSTD have more chances to be reproduced in independent studies compared to the less stable ones. At the same time, we show that a transcriptomics dataset can be reduced to a relatively high number of dimensions without losing the interpretability of ICA, even though higher dimensions give rise to components driven by small gene sets. We suggest a protocol of ICA application to transcriptomics data with a possibility of prioritizing components with respect to their reproducibility that strengthens the biological interpretation. Computing too few components (much less than MSTD) is not optimal for interpretability of the results. The components ranked within MSTD range have more chances to be reproduced in independent studies. The online version of this article (10.1186/s12864-017-4112-9) contains supplementary material, which is available to authorized users.
DOI: 10.1038/ng.2764
发表时间: 2013-10
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Weinstein, John N.;Collisson, Eric A.;Mills, Gordon B.;Shaw, Kenna R. Mills;Ozenberger, Brad A.;Ellrott, Kyle;Shmulevich, Ilya;Sander, Chris;Stuart, Joshua M.
通讯作者: Stuart, Joshua M.
DOI: 10.1016/s0893-6080(00)00026-5
发表时间: 2000-05-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
Hyvärinen, A;Oja, E
通讯作者: Oja, E
DOI: 10.1038/sj.onc.1207562
发表时间: 2004-08-26
期刊: ONCOGENE
影响因子: 8
作者:
Saidi, SA;Holland, CM;Smith, SK
通讯作者: Smith, SK
DOI: 10.1093/nar/gkp427
发表时间: 2009-07
影响因子: 14.9
作者:
Chen J;Bardes EE;Aronow BJ;Jegga AG
通讯作者: Jegga AG
DOI: 10.1186/s12864-016-3435-2
发表时间: 2017-01-05
期刊: BMC genomics
影响因子: 4.4
作者:
Giotti B;Joshi A;Freeman TC
通讯作者: Freeman TC