Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samples

Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samples
复制标题

DOI:
10.1093/bib/bbab265
复制
发表时间:
2021-08-04
影响因子:
9.5
通讯作者:
Mangul, Serghei
Mangul, Serghei
中科院分区:
生物学2区
文献类型:
--
作者:
Nadel, Brian B.;Oliva, Meritxell;Mangul, Serghei

文献摘要

被引文献

相似文献

估计血液和组织样品的细胞类型组成是实验室研究和临床护理中相关的生物学挑战。近年来,已经开发了许多计算工具来使用基因表达数据估计细胞类型的丰度。尽管这些工具使用了多种方法,但它们都利用纯化的细胞类型的表达曲线来评估样品中的细胞类型组成。在这项研究中,我们比较了12种细胞类型的定量工具,并在使用10个单独的参考曲线中的每一个中评估其性能。具体来说,我们已经在4000多个样品上运行了每个工具,具有已知的细胞类型比例,涵盖了免疫细胞类型和基质细胞类型。其中总共有12个代表了体外合成混合物和300个使用单细胞数据制备的硅合成混合物中。最终的3728个临床样品是从弗雷明汉队列中收集的,使用电阻抗细胞计数对细胞群进行了定量。当将工具应用于Framingham数据集时,估计免疫细胞和癌细胞比例(EPIC)的工具会产生最高的相关性,而基因表达反卷积互动工具(GEDIT)产生的误差最低。其他数据集的最佳工具是多种多样的,但是Cibersort和GEDIT最一致地产生了准确的结果。我们发现,最佳参考取决于所使用的工具,并报告建议与每个工具一起使用的参考。大多数工具在几分钟内返回结果,但是在大型数据集中,Cibersort的运行时间可能超过小时甚至几天。我们得出的结论是,反卷积方法能够返回高质量的结果,但是正确的参考选择至关重要。
Estimating cell type composition of blood and tissue samples is a biological challenge relevant in both laboratory studies and clinical care. In recent years, a number of computational tools have been developed to estimate cell type abundance using gene expression data. Although these tools use a variety of approaches, they all leverage expression profiles from purified cell types to evaluate the cell type composition within samples. In this study, we compare 12 cell type quantification tools and evaluate their performance while using each of 10 separate reference profiles. Specifically, we have run each tool on over 4000 samples with known cell type proportions, spanning both immune and stromal cell types. A total of 12 of these represent in vitro synthetic mixtures and 300 represent in silico synthetic mixtures prepared using single-cell data. A final 3728 clinical samples have been collected from the Framingham cohort, for which cell populations have been quantified using electrical impedance cell counting. When tools are applied to the Framingham dataset, the tool Estimating the Proportions of Immune and Cancer cells (EPIC) produces the highest correlation, whereas Gene Expression Deconvolution Interactive Tool (GEDIT) produces the lowest error. The best tool for other datasets is varied, but CIBERSORT and GEDIT most consistently produce accurate results. We find that optimal reference depends on the tool used, and report suggested references to be used with each tool. Most tools return results within minutes, but on large datasets runtimes for CIBERSORT can exceed hours or even days. We conclude that deconvolution methods are capable of returning high-quality results, but that proper reference selection is critical.