High-Order Correlation Integration for Single-Cell or Bulk RNA-seq Data Analysis

High-Order Correlation Integration for Single-Cell or Bulk RNA-seq Data Analysis
复制标题

用于单细胞或批量 RNA-seq 数据分析的高阶相关积分

DOI:
10.3389/fgene.2019.00371
复制
发表时间:
2019-04
影响因子:
3.7
通讯作者:
Luonan Chen
Luonan Chen
中科院分区:
生物学3区
文献类型:
--
作者:
Hui Tang;Tao Zeng;Luonan Chen

文献摘要

参考文献

被引文献

相似文献

高质量地量化或标记样本类型是一项具有挑战性的任务,这是理解复杂疾病的关键一步。减少数据的噪声污染并确保提取的内在模式与原始数据结构一致对于样本聚类和分类非常重要。在这里,我们提出了一种有效的数据集成框架,称为HCI(高阶相关集成),它利用高阶相关矩阵与模式融合分析(PFA)相结合,实现高维数据特征提取。一方面,高阶皮尔逊相关系数可以突出噪声输入数据集背后的潜在模式,从而提高当前可用于样本聚类的算法的准确性和鲁棒性。另一方面,PFA 可以通过优化调整信号效应,从不同的输入矩阵中有效地识别内在样本模式。为了验证我们新方法的有效性,我们首先在四个单细胞RNA-seq数据集上应用HCI来区分细胞类型,我们发现HCI能够从scRNA-seq数据中识别单细胞样本的先前已知细胞类型,在不同条件下比其他方法具有更高的准确性和鲁棒性。其次,我们还整合了来自 TCGA 数据集和 GEO 数据集的异质组学数据,包括批量 RNA-seq 数据,该数据在识别不同癌症亚型方面优于其他方法。在另一个案例研究中,我们还根据 HCI 估计的特征权重构建了结直肠癌的 mRNA-miRNA 调控网络,其中差异表达的 mRNA 和 miRNA 在众所周知的结直肠癌功能集中显着富集,例如 KEGG 通路和 IPA 疾病注释。所有这些结果都表明HCI对于不同类型和组织的RNA-seq数据的样本聚类具有广泛的灵活性和适用性。
Quantifying or labeling the sample type with high quality is a challenging task, which is a key step for understanding complex diseases. Reducing noise pollution to data and ensuring the extracted intrinsic patterns in concordance with the primary data structure are important in sample clustering and classification. Here we propose an effective data integration framework named as HCI (High-order Correlation Integration), which takes an advantage of high-order correlation matrix incorporated with pattern fusion analysis (PFA), to realize high-dimensional data feature extraction. On the one hand, the high-order Pearson's correlation coefficient can highlight the latent patterns underlying noisy input datasets and thus improve the accuracy and robustness of the algorithms currently available for sample clustering. On the other hand, the PFA can identify intrinsic sample patterns efficiently from different input matrices by optimally adjusting the signal effects. To validate the effectiveness of our new method, we firstly applied HCI on four single-cell RNA-seq datasets to distinguish the cell types, and we found that HCI is capable of identifying the prior-known cell types of single-cell samples from scRNA-seq data with higher accuracy and robustness than other methods under different conditions. Secondly, we also integrated heterogonous omics data from TCGA datasets and GEO datasets including bulk RNA-seq data, which outperformed the other methods at identifying distinct cancer subtypes. Within an additional case study, we also constructed the mRNA-miRNA regulatory network of colorectal cancer based on the feature weight estimated from HCI, where the differentially expressed mRNAs and miRNAs were significantly enriched in well-known functional sets of colorectal cancer, such as KEGG pathways and IPA disease annotations. All these results supported that HCI has extensive flexibility and applicability on sample clustering with different types and organizations of RNA-seq data.
一种寻找目标控制复杂网络最佳驱动节点的新算法及其在药物靶标识别中的应用
DOI: 10.1186/s12864-017-4332-z
发表时间: 2018-01-19
期刊: BMC genomics
影响因子: 4.4
作者:
Guo WF;Zhang SW;Shi QQ;Zhang CM;Zeng T;Chen L
通讯作者: Chen L
使用样本特定网络对疾病进行个性化表征
DOI: 10.1093/nar/gkw772
发表时间: 2016-12-15
影响因子: 14.9
作者:
Liu X;Wang Y;Ji H;Aihara K;Chen L
通讯作者: Chen L
DOI: 10.1093/bioinformatics/btr206
发表时间: 2011-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Zhang S;Li Q;Liu J;Zhou XJ
通讯作者: Zhou XJ
DOI: 10.1038/s41467-018-03024-2
发表时间: 2018-02-14
影响因子: 16.6
作者:
Yang B;Li M;Tang W;Liu W;Zhang S;Chen L;Xia J
通讯作者: Xia J
单细胞RNA测序揭示了T辅助细胞合成从头开始的类固醇,从而有助于免疫稳态。
DOI: 10.1016/j.celrep.2014.04.011
发表时间: 2014-05-22
期刊: Cell reports
影响因子: 8.8
作者:
Mahata B;Zhang X;Kolodziejczyk AA;Proserpio V;Haim-Vilmovsky L;Taylor AE;Hebenstreit D;Dingler FA;Moignard V;Göttgens B;Arlt W;McKenzie AN;Teichmann SA
通讯作者: Teichmann SA