Leveraging heterogeneity across multiple datasets increases cell-mixture deconvolution accuracy and reduces biological and technical biases.

Leveraging heterogeneity across multiple datasets increases cell-mixture deconvolution accuracy and reduces biological and technical biases.
复制标题

DOI:
10.1038/s41467-018-07242-6
复制
发表时间:
2018-11-09
影响因子:
16.6
通讯作者:
Khatri P
Khatri P
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Vallania F;Tam A;Lofgren S;Schaffert S;Azad TD;Bongen E;Haynes W;Alsup M;Alonso M;Davis M;Engleman E;Khatri P

文献摘要

参考文献

被引文献

相似文献

来自混合细胞转录组学数据的细胞比例的计算机定量(去卷积)需要参考表达矩阵,称为基础矩阵。我们假设仅使用来自单个微阵列平台的健康样本创建的矩阵将在去卷积中引入生物和技术偏差。我们在两个现有的矩阵,IRIS和LM 22,无论反卷积方法存在这样的偏见。在这里,我们提出了immunoStates,一个基础矩阵,使用42个微阵列平台上的6160个不同疾病状态的样本构建。我们发现,immunoStates显著降低了生物和技术偏见。重要的是,我们发现,一旦选择了基矩阵,不同的方法几乎没有影响或影响很小。我们进一步表明,在所有方法中,使用免疫状态的细胞比例估计值与测量比例的相关性始终高于IRIS和LM 22。我们的研究结果表明,将生物和技术异质性的基础矩阵,以实现一贯的高精度的需要和重要性。来自批量表达数据的细胞类型去卷积依赖于参考表达矩阵。在这里,作者介绍了一个基础矩阵,该矩阵使用来自42个平台上的健康和患病样本的数据构建,减少了使用健康样本构建的单平台矩阵引入的偏差。
In silico quantification of cell proportions from mixed-cell transcriptomics data (deconvolution) requires a reference expression matrix, called basis matrix. We hypothesize that matrices created using only healthy samples from a single microarray platform would introduce biological and technical biases in deconvolution. We show presence of such biases in two existing matrices, IRIS and LM22, irrespective of deconvolution method. Here, we present immunoStates, a basis matrix built using 6160 samples with different disease states across 42 microarray platforms. We find that immunoStates significantly reduces biological and technical biases. Importantly, we find that different methods have virtually no or minimal effect once the basis matrix is chosen. We further show that cellular proportion estimates using immunoStates are consistently more correlated with measured proportions than IRIS and LM22, across all methods. Our results demonstrate the need and importance of incorporating biological and technical heterogeneity in a basis matrix for achieving consistently high accuracy. Cell type deconvolution from bulk expression data rely on a reference expression matrix. Here, the authors introduce a basis matrix built using data from both healthy and diseased samples profiled on 42 platforms, reducing biases introduced by single-platform matrices built using healthy samples.
全基因组诊断肺结核的表达:一种多螺旋分析。
DOI: 10.1016/s2213-2600(16)00048-5
发表时间: 2016-03
期刊: The Lancet. Respiratory medicine
影响因子: --
作者:
Sweeney TE;Braviak L;Tato CM;Khatri P
通讯作者: Khatri P
DOI: 10.1126/scitranslmed.aaf7165
发表时间: 2016-07-06
影响因子: 17.1
作者:
Sweeney TE;Wong HR;Khatri P
通讯作者: Khatri P
DOI: 10.1126/scitranslmed.aaa5993
发表时间: 2015-05-13
影响因子: 17.1
作者:
Sweeney TE;Shidham A;Wong HR;Khatri P
通讯作者: Khatri P
DOI: 10.1016/j.coi.2013.09.015
发表时间: 2013-10
影响因子: 7
作者:
Shen-Orr SS;Gaujoux R
通讯作者: Gaujoux R
DOI: 10.1038/nature13320
发表时间: 2014-06-12
期刊: NATURE
影响因子: 64.8
作者:
Mazur, Pawel K.;Reynoird, Nicolas;Khatri, Purvesh;Jansen, Pascal W. T. C.;Wilkinson, Alex W.;Liu, Shichong;Barbash, Olena;Van Aller, Glenn S.;Huddleston, Michael;Dhanak, Dashyant;Tummino, Peter J.;Kruger, Ryan G.;Garcia, Benjamin A.;Butte, Atul J.;Vermeulen, Michiel;Sage, Julien;Gozani, Or
通讯作者: Gozani, Or