SHINE: protein language model-based pathogenicity prediction for short inframe insertion and deletion variants.

SHINE: protein language model-based pathogenicity prediction for short inframe insertion and deletion variants.
复制标题

SHINE:基于蛋白质语言模型的短内框插入和缺失变异的致病性预测。

DOI:
10.1093/bib/bbac584
复制
发表时间:
2023
影响因子:
9.5
通讯作者:
Shen,Yufeng
Shen,Yufeng
中科院分区:
生物学2区
文献类型:
--
作者:
Fan,Xiao;Pan,Hongbing;Tian,Alan;Chung,WendyK;Shen,Yufeng

文献摘要

相似文献

准确的变异致病性预测在人类疾病的遗传学研究中非常重要。框内插入和缺失变异体(INDELs)改变蛋白质序列和长度,但不像移码INDELs那样有害。由于可用于培训的已知致病变异体的数量有限,INFRAME INDELL解释具有挑战性。现有的预测方法大多使用人工编码的特征,包括保守性、蛋白质结构和功能以及等位基因频率来推断变异致病。蛋白质序列和结构的深度学习建模的最新进展为改进基于大量蛋白质序列的显著特征的表示提供了机会。我们开发了一种新的SHortInFrame插入和缺失(SISH)致病预测因子。Share使用预先训练的蛋白质语言模型从蛋白质序列和多个蛋白质序列比对中构建Indel及其蛋白质上下文的潜在表示,并将该潜在表示馈送到有监督的机器学习模型中进行致病性预测。我们整理了来自ClinVar和gnomAD的训练数据,并从不同的来源创建了两个测试数据集。对于这两个测试数据集中的缺失和插入变体,SISH获得了比现有方法更好的预测性能。我们的工作表明,非监督蛋白质语言模型可以提供关于蛋白质的有价值的信息,基于这些模型的新方法可以改进遗传分析中的变体解释。
Accurate variant pathogenicity predictions are important in genetic studies of human diseases. Inframe insertion and deletion variants (indels) alter protein sequence and length, but not as deleterious as frameshift indels. Inframe indel Interpretation is challenging due to limitations in the available number of known pathogenic variants for training. Existing prediction methods largely use manually encoded features including conservation, protein structure and function, and allele frequency to infer variant pathogenicity. Recent advances in deep learning modeling of protein sequences and structures provide an opportunity to improve the representation of salient features based on large numbers of protein sequences. We developed a new pathogenicity predictor forSHortInframe iNsertion and dEletion (SHINE). SHINE uses pretrained protein language models to construct a latent representation of an indel and its protein context from protein sequences and multiple protein sequence alignments, and feeds the latent representation into supervised machine learning models for pathogenicity prediction. We curated training data from ClinVar and gnomAD, and created two test datasets from different sources. SHINE achieved better prediction performance than existing methods for both deletion and insertion variants in these two test datasets. Our work suggests that unsupervised protein language models can provide valuable information about proteins, and new methods based on these models can improve variant interpretation in genetic analyses.