Building trans-omics evidence: using imaging and 'omics' to characterize cancer profiles

Building trans-omics evidence: using imaging and 'omics' to characterize cancer profiles
复制标题

建立跨组学证据:使用成像和“组学”来表征癌症特征

DOI:
--
复制
发表时间:
2018
期刊:
Pacific Symposium on Biocomputing
影响因子:
--
通讯作者:
R. Machiraju
R. Machiraju
中科院分区:
--
文献类型:
--
作者:
Arunima Srivastava;Chaitanya Kulkarni;P. Mallick;Kun Huang;R. Machiraju

文献摘要

相似文献

利用单一模式数据在癌症中建立预测模型导致了对大多数患者概况的相当狭窄的看法。S的一些临床方面与组织学影像特征密切相关,例如肿瘤分期,而另一些方面与基因组和蛋白质组变异(例如癌症亚型和疾病侵袭性生物标志物)有关。我们假设存在连贯的“跨组学”特征,这些特征描述了多个数据来源的不同临床队列,导致了更具描述性和健壮性的疾病特征。在这项工作中,对于来自TCGA(癌症基因组图谱)的L 05乳腺癌患者,我们考虑了四个临床属性(AJCC分期、肿瘤分期、ER状态和PAM50mRNA亚型),并使用三种不同的数据形式(组织病理学图像、转录组学和蛋白质组学)建立预测模型。随后,我们确定了关键的多层次特征,这些特征推动了对不同不同队列的患者的成功分类。为了建立每种数据类型的预测值,我们采用了广泛使用的“最佳实践”技术,包括用于组织病理学图像的基于CNN(卷积神经网络)的分类器和用于蛋白质组数据的回归模型。正如预期的那样,组织学图像在预测癌症分期方面优于分子特征,转录学对ER状态和PAM50亚型具有优越的区分力,但在少数情况下,所有数据模式的表现都具有可比性。此外,我们还确定了一组关键基因和蛋白质,它们的表达和丰度在每个临床队列中都相关,包括(I)肿瘤的严重程度和进展(包括。GABARAP)、(Ii)ER状态(包括ESR1)和(Iii)疾病亚型(包括FOXCL)。因此,我们定量评估了不同数据类型在预测关键乳腺癌患者属性和改善疾病特征方面的有效性。
Utilization of single modality data to build predictive models in cancer results in a rather narrow view of most patient profiles. Some clinical facet s relate strongly to histology image features, e.g. tumor stages, whereas others are associated with genomic and proteomic variations (e.g. cancer subtypes and disease aggression biomarkers). We hypothesize that there are coherent "trans-omics" features that characterize varied clinical cohorts across multiple sources of data leading to more descriptive and robust disease characterization. In this work, for l 05 breast cancer patients from the TCGA (The Cancer Genome Atlas), we consider four clinical attributes (AJCC Stage, Tumor Stage, ER-Status and PAM50 mRNA Subtypes), and build predictive models using three different modalities of data (histopathological images, transcriptomics and proteomics). Following which, we identify critical multi-level features that drive successful classification of patients for the various different cohorts. To build predictors for each data type, we employ widely used "best practice" techniques including CNN-based (convolutional neural network) classifiers for histopathological images and regression models for proteogenomic data. While, as expected, histology images outperformed molecular features while predicting cancer stages, and transcriptomics held superior discriminatory power for ER-Status and PAM50 subtypes, there exist a few cases where all data modalities exhibited comparable performance. Further, we also identified sets of key genes and proteins whose expression and abundance correlate across each clinical cohort including (i) tumor severity and progression (incl. GABARAP), (ii) ER-status (incl.ESRl) and (iii) disease subtypes (incl. FOXCl). Thus, we quantitatively assess the efficacy of different data types to predict critical breast cancer patient attributes and improve disease characterization.