Deep Learning to Discover Cancer Glycome Genes Signifying the Origins of Cancer

Deep Learning to Discover Cancer Glycome Genes Signifying the Origins of Cancer
复制标题

DOI:
10.1109/bibm49941.2020.9313450
复制
发表时间:
2020-12
期刊:
2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)
影响因子:
--
通讯作者:
Abdullah Al Mamun;Masrur Sobhan;R. Tanvir;C. Dimitroff;A. Mondal
Abdullah Al Mamun;Masrur Sobhan;R. Tanvir;C. Dimitroff;A. Mondal
中科院分区:
其他
文献类型:
--
作者:
Abdullah Al Mamun;Masrur Sobhan;R. Tanvir;C. Dimitroff;A. Mondal

文献摘要

被引文献

相似文献

背景:蛋白质糖基化异常是肿瘤的一个常见特征,并与肿瘤的恶性行为有关。然而,细胞糖组如何以及在多大程度上参与癌症的发展和进展仍然不确定。本研究的主要目的是利用癌症基因组的表达谱对糖组基因进行计算机识别,从而揭示癌症的特征。有一个列表中的$\sim 500$糖组基因在几个分子类别。本研究基于以下假设:如果糖基化是癌症的共同特征,则存在癌症糖组基因的短名单,并且它们的表达谱应携带能够区分癌症基因组图谱(TCGA)中可用的33种不同癌症的特征。TCGA中癌症样本的分布高度不平衡,范围从胆管癌(CHOL)的36例到乳腺癌(BRCA)的1089例。用于识别特征基因的监督特征选择方法将偏向于较大的群体。我们开发了一个使用混凝土自动编码器(CAE)的计算框架,这是一种基于深度学习的无监督特征选择算法,可以找到癌症相关的糖组基因。本研究中使用的最佳特征子集的标准是(a)特征的数量应尽可能少,(B)使用所选特征的分类准确率应> 90%。结果:我们的实验显示了一个糖组基因的候选名单(132个基因),可以区分33种不同的癌症,准确率为92%。这一研究表明,癌症糖组基因标志着癌症的起源。
Background: Aberrant protein glycosylation is a common feature of cancer and contributes to malignant behavior. However, how and to what extent the cellular glycome is involved in cancer development and progression is still undefined. The primary objective of this study is to conduct insilico identification of glycome genes that could reveal a signature of cancer using expression profiles of cancer genomes. There exists a list of $\sim 500$ glycome genes in several molecular categories. This study is based on the hypothesis that if the glycosylation is a common feature of cancer, there exists a shortlist of cancer glycome genes and their expression profiles should carry the signature capable of differentiating 33 different cancers available in The Cancer Genome Atlas (TCGA).Method: The distribution of cancer samples in TCGA is highly imbalanced, ranging from 36 for Cholangiocarcinoma (CHOL) to 1089 for Breast Cancer (BRCA). Supervised feature selection approaches to identify the signature genes would be biased to larger groups. We developed a computational framework using concrete autoencoder (CAE), a deep learning-based unsupervised feature selection algorithm, to find the cancer-related glycome genes. The criteria of optimal feature subset used in this study are (a) the number of features should be as few as possible, and (b) accuracy of classification using the selected features should be >90%.Results: Our experiment showed a shortlist of glycome genes (132 genes) that can differentiate 33 different cancers with an accuracy of 92%. This study reflects that the cancer glycome genes signify the origins of cancer.