Cancer Subtyping via Embedded Unsupervised Learning on Transcriptomics Data

Cancer Subtyping via Embedded Unsupervised Learning on Transcriptomics Data
复制标题

DOI:
10.1109/embc48229.2022.9870903
复制
发表时间:
2022-04
期刊:
2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC)
影响因子:
--
通讯作者:
Ziwei Yang;Lingwei Zhu;Zheng Chen;Ming Huang;N. Ono;M. Altaf-Ul-Amin;S. Kanaya
Ziwei Yang;Lingwei Zhu;Zheng Chen;Ming Huang;N. Ono;M. Altaf-Ul-Amin;S. Kanaya
中科院分区:
其他
文献类型:
--
作者:
Ziwei Yang;Lingwei Zhu;Zheng Chen;Ming Huang;N. Ono;M. Altaf-Ul-Amin;S. Kanaya

文献摘要

相似文献

癌症是世界上最致命的疾病之一。肿瘤亚型的准确诊断和分类对于有效的临床治疗是必不可少的。随着各种深度学习方法的出现,最近发表了关于自动癌症分型系统的有希望的结果。然而,这样的自动系统往往过拟合的数据,由于高维和稀缺性。在本文中,我们建议从无监督学习的角度来研究自动子类型,直接构建底层数据分布本身,因此可以生成足够的数据来缓解过拟合问题。具体来说,我们绕过了强高斯性假设,该假设通常存在,但由于矢量量化的小样本而在无监督学习子类型文献中失败。正如大量实验结果所证明的那样,我们提出的方法更好地捕获了潜在的空间特征,并在分子基础上对癌症亚型表现进行了建模。
Cancer is one of the deadliest diseases worldwide. Accurate diagnosis and classification of cancer subtypes are indispensable for effective clinical treatment. Promising results on automatic cancer subtyping systems have been published recently with the emergence of various deep learning methods. However, such automatic systems often overfit the data due to the high dimensionality and scarcity. In this paper, we propose to investigate automatic subtyping from an unsupervised learning perspective by directly constructing the underlying data distribution itself, hence sufficient data can be generated to alleviate the issue of overfitting. Specifically, we bypass the strong Gaussianity assumption that typically exists but fails in the unsupervised learning subtyping literature due to small-sized samples by vector quantization. Our proposed method better captures the latent space features and models the cancer subtype manifestation on a molecular basis, as demonstrated by the extensive experimental results.