Correlation-centred variable selection of a gene expression signature to predict breast cancer metastasis

Correlation-centred variable selection of a gene expression signature to predict breast cancer metastasis
复制标题

DOI:
10.1038/s41598-020-64870-z
复制
发表时间:
2020-05-13
期刊:
影响因子:
4.6
通讯作者:
Tomita, Masaru
Tomita, Masaru
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hikichi, Shiori;Sugimoto, Masahiro;Tomita, Masaru

文献摘要

被引文献

相似文献

人们深入研究基于基因特征的远处癌症转移预测,以实现精确的诊断和治疗。基因选择,即特征选择,是建立准确预测和理解潜在病理学的基石。在这里,我们开发了一种简单但强大的特征选择方法,使用以相关性为中心的方法来选择具有高预测和泛化能力的最小基因集。使用多元逻辑回归模型来预测乳腺癌患者的 5 年转移情况。从淋巴结阴性乳腺癌患者的肿瘤样本中获得的基因表达数据被随机分为训练数据和验证数据。我们的方法使用训练数据选择了 12 个基因,与之前报道的 76 个基因产生的 0.579 相比,这显示了 0.730 的接收者操作特征曲线下面积更高。预测模型的签名在独立数据集中得到验证,并观察到其更高的泛化能力。基因本体分析表明,我们的方法一致地选择了 76 个基因经常选择的具有相同功能的基因。总而言之,我们的方法识别出具有高预测能力的较少基因组,这些基因组具有多种用途并适用于预测其他因素,例如药物治疗的结果和其他癌症类型的预后。
Predictions of distant cancer metastasis based on gene signatures are studied intensively to realise precise diagnosis and treatments. Gene selection i.e. feature selection is a cornerstone to both establish accurate predictions and understand underlying pathologies. Here, we developed a simple but robust feature selection method using a correlation-centred approach to select minimal gene sets that have both high predictive and generalisation abilities. A multiple logistic regression model was used to predict 5-year metastases of patients with breast cancer. Gene expression data obtained from tumour samples of lymph node-negative breast cancer patients were randomly split into training and validation data. Our method selected 12 genes using training data and this showed a higher area under the receiver operating characteristic curve of 0.730 compared with 0.579 yielded by previously reported 76 genes. The signature with the predictive model was validated in an independent dataset, and its higher generalization ability was observed. Gene ontology analyses revealed that our method consistently selected genes with identical functions which frequently selected by the 76 genes. Taken together, our method identifies fewer gene sets bearing high predictive abilities, which would be versatile and applicable to predict other factors such as the outcomes of medical treatments and prognoses of other cancer types.