Exploring the application of deep learning methods for polygenic risk score estimation

Exploring the application of deep learning methods for polygenic risk score estimation
复制标题

探索深度学习方法在多基因风险评分估计中的应用

DOI:
10.1101/2023.12.14.23299972
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Squires S
Squires S
中科院分区:
--
文献类型:
--
作者:
Squires S

文献摘要

参考文献

相似文献

背景多基因风险评分(PR)将遗传信息总结为一个具有临床和研究用途的数字。深度学习已经给多个领域带来了革命性的变化,然而,深度学习对PRSS的影响并不那么显著。我们探索了动态链接库如何改进PR的生成。方法我们使用英国生物库数据在已知PRS上训练动态链接库模型。我们探讨了这些模型是否能够重建人类编程的PR,包括使用一个模型来生成多个PR,以及在PR生成中的DL困难。我们研究了动态学习如何弥补缺失数据和对性能的限制。结果我们展示了在减少训练数据量的情况下,在几乎没有性能损失的情况下,几乎完美地生成多个PRS。对于一组丢失的SNP,DL模型产生的预测能够将病例从总体样本中分离出来,接收器操作特征曲线下的面积为0.847(95%CI:0.828-0.864),而PR的面积为0.798(95%CI:0.779-0.818)。结论DL可以准确地生成PRS,包括一个模型用于多个PRS。该模型具有可移植和高寿命的特点。对于某些缺失的SNP,DL模型可以改进PR的生成;进一步的改进可能需要额外的输入数据。
BackgroundPolygenic risk scores (PRS) summarise genetic information into a single number with clinical and research uses. Deep learning (DL) has revolutionised multiple fields, however, the impact of DL on PRSs has been less significant. We explore how DL can improve the generation of PRSs.MethodsWe train DL models on known PRSs using UK Biobank data. We explore whether the models can recreate human programmed PRSs, including using a single model to generate multiple PRSs, and DL difficulties in PRS generation. We investigate how DL can compensate for missing data and constraints on performance.ResultsWe demonstrate almost perfect generation of multiple PRSs with little loss of performance with reduced quantity of training data. For an example set of missing SNPs the DL model produces predictions that enable separation of cases from population samples with an area under the receiver operating characteristic curve of 0.847 (95% CI: 0.828–0.864) compared to 0.798 (95% CI: 0.779–0.818) for the PRS.ConclusionsDL can accurately generate PRSs, including with one model for multiple PRSs. The models are transferable and have high longevity. With certain missing SNPs the DL models can improve on PRS generation; further improvements would likely require additional input data.
VIME:将自监督和半监督学习的成功扩展到表格领域
DOI: --
发表时间: 2020
期刊: Neural Information Processing Systems
影响因子: --
作者:
Jinsung Yoon;Yao Zhang;James Jordon;M. Schaar
通讯作者: M. Schaar