Characterizing and explaining impact of disease-associated mutations in proteins without known structures or structural homologues

Characterizing and explaining impact of disease-associated mutations in proteins without known structures or structural homologues
复制标题

DOI:
10.1101/2021.11.17.468998
复制
发表时间:
2021-11
期刊:
bioRxiv
影响因子:
--
通讯作者:
Neeladri Sen;I. Anishchenko;N. Bordin;I. Sillitoe;S. Velankar;D. Baker;C. Orengo
Neeladri Sen;I. Anishchenko;N. Bordin;I. Sillitoe;S. Velankar;D. Baker;C. Orengo
中科院分区:
其他
文献类型:
--
作者:
Neeladri Sen;I. Anishchenko;N. Bordin;I. Sillitoe;S. Velankar;D. Baker;C. Orengo

文献摘要

相似文献

人类蛋白质的突变会导致疾病。这些蛋白质的结构有助于了解这些疾病的机制,并开发针对它们的治疗方法。利用改进的深度学习技术,如RoseTTAFold和AlphaFold,即使在没有结构同源的情况下,我们也可以预测蛋白质的结构。我们从蛋白质数据库(PDB)中没有已知蛋白质结构或密切同源的553个与疾病相关的人类蛋白质中建模并提取了结构域。我们注意到,与只能分配到结构未知的Pfam家族或无法分配到这两个家族的结构域相比,AlphaFold和RoseTTAFold模型之间的模型质量更高,RMSD更低。我们预测了这些预测结构中的配体结合位点、蛋白质-蛋白质界面和保守残基。然后,我们探索了与疾病相关的错义突变是否位于这些预测的功能位点的附近,是否根据DDG计算破坏了蛋白质结构的稳定,或者是否预测它们是致病的。我们可以解释80%的这些与疾病相关的突变是基于与功能位点的接近、结构不稳定或致病性。与多态相比,与疾病相关的错义突变被掩埋的比例更大,更接近预测的功能位点,预测为不稳定和/或致病。使用两种最先进的技术中的模型为我们的预测提供了更好的信心,我们解释了基于RoseTAFold模型的93个额外的突变,这些突变不能仅基于AlphaFold模型来解释。
Mutations in human proteins lead to diseases. The structure of these proteins can help understand the mechanism of such diseases and develop therapeutics against them. With improved deep learning techniques such as RoseTTAFold and AlphaFold, we can predict the structure of proteins even in the absence of structural homologues. We modeled and extracted the domains from 553 disease-associated human proteins without known protein structures or close homologues in the Protein Databank (PDB). We noticed that the model quality was higher and the RMSD lower between AlphaFold and RoseTTAFold models for domains that could be assigned to CATH families as compared to those which could only be assigned to Pfam families of unknown structure or could not be assigned to either. We predicted ligand-binding sites, protein-protein interfaces, conserved residues in these predicted structures. We then explored whether the disease-associated missense mutations were in the proximity of these predicted functional sites, if they destabilized the protein structure based on ddG calculations or if they were predicted to be pathogenic. We could explain 80% of these disease-associated mutations based on proximity to functional sites, structural destabilization or pathogenicity. When compared to polymorphisms a larger percentage of disease associated missense mutations were buried, closer to predicted functional sites, predicted as destabilising and/or pathogenic. Usage of models from the two state-of-the-art techniques provide better confidence in our predictions, and we explain 93 additional mutations based on RoseTTAFold models which could not be explained based solely on AlphaFold models.