Multiple-instance learning of somatic mutations for the classification of tumour type and the prediction of microsatellite status.

Multiple-instance learning of somatic mutations for the classification of tumour type and the prediction of microsatellite status.
复制标题

体细胞突变的多实例学习用于肿瘤类型分类和微卫星状态预测。

DOI:
10.1038/s41551-023-01120-3
复制
发表时间:
2024-01
影响因子:
28.1
通讯作者:
--
中科院分区:
工程技术1区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

大规模基因组数据非常适合通过深度学习算法进行分析。然而,对于许多基因组数据集,标签是在样本的水平上,而不是针对单个基因组测量。利用这些数据集的机器学习模型通过使用静态编码的度量来生成预测,然后在样本级别进行聚合。在这里,我们证明了一个具有多头注意力的单一弱监督端到端多实例学习模型可以被训练来编码和聚合体细胞突变的局部序列背景或基因组位置,从而允许对样本级分类的个体测量的重要性进行建模,从而提供增强的可解释性。该模型解决了传统模型无法完成的综合任务,并在肿瘤类型分类和预测微卫星状态方面实现了同类最佳性能。通过提高需要来自基因组数据集的聚合信息的任务的性能,多实例深度学习可以生成生物学洞察力。一个多实例学习模型被训练来编码和聚合局部序列背景或体细胞突变的基因组位置,在分类和预测任务中取得了同类最佳的性能。
Large-scale genomic data are well suited to analysis by deep learning algorithms. However, for many genomic datasets, labels are at the level of the sample rather than for individual genomic measures. Machine learning models leveraging these datasets generate predictions by using statically encoded measures that are then aggregated at the sample level. Here we show that a single weakly supervised end-to-end multiple-instance-learning model with multi-headed attention can be trained to encode and aggregate the local sequence context or genomic position of somatic mutations, hence allowing for the modelling of the importance of individual measures for sample-level classification and thus providing enhanced explainability. The model solves synthetic tasks that conventional models fail at, and achieves best-in-class performance for the classification of tumour type and for predicting microsatellite status. By improving the performance of tasks that require aggregate information from genomic datasets, multiple-instance deep learning may generate biological insight. A multiple-instance-learning model trained to encode and aggregate either the local sequence contexts or the genomic positions of somatic mutations achieved best-in-class performance in classification and prediction tasks.
DOI: 10.1038/s41551-020-00682-w
发表时间: 2021-06
影响因子: 28.1
作者:
Lu MY;Williamson DFK;Chen TY;Chen RJ;Barbieri M;Mahmood F
通讯作者: Mahmood F
DOI: 10.1038/nature12113
发表时间: 2013-05-02
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1016/j.ccell.2018.03.014
发表时间: 2018-04-09
期刊: Cancer cell
影响因子: 50.3
作者:
Berger AC;Korkut A;Kanchi RS;Hegde AM;Lenoir W;Liu W;Liu Y;Fan H;Shen H;Ravikumar V;Rao A;Schultz A;Li X;Sumazin P;Williams C;Mestdagh P;Gunaratne PH;Yau C;Bowlby R;Robertson AG;Tiezzi DG;Wang C;Cherniack AD;Godwin AK;Kuderer NM;Rader JS;Zuna RE;Sood AK;Lazar AJ;Ojesina AI;Adebamowo C;Adebamowo SN;Baggerly KA;Chen TW;Chiu HS;Lefever S;Liu L;MacKenzie K;Orsulic S;Roszik J;Shelley CS;Song Q;Vellano CP;Wentzensen N;Cancer Genome Atlas Research Network;Weinstein JN;Mills GB;Levine DA;Akbani R
通讯作者: Akbani R
DOI: 10.1200/po.17.00073
发表时间: 2017
影响因子: 4.6
作者:
Bonneville R;Krook MA;Kautto EA;Miya J;Wing MR;Chen HZ;Reeser JW;Yu L;Roychowdhury S
通讯作者: Roychowdhury S
DOI: 10.1016/j.artint.2013.06.003
发表时间: 2013-08-01
影响因子: 14.4
作者:
Amores, Jaume
通讯作者: Amores, Jaume