Deep generative modeling of the human proteome reveals over a hundred novel genes involved in rare genetic disorders.

Deep generative modeling of the human proteome reveals over a hundred novel genes involved in rare genetic disorders.
复制标题

人类蛋白质组的深度生成模型揭示了一百多个与罕见遗传疾病有关的新基因。

DOI:
10.1101/2023.11.27.23299062
复制
发表时间:
2023
期刊:
medRxiv : the preprint server for health sciences
影响因子:
--
通讯作者:
Marks,DeboraS
Marks,DeboraS
中科院分区:
--
文献类型:
--
作者:
Orenbuch,Rose;Kollasch,AaronW;Spinner,HansenD;Shearer,CourtneyA;Hopf,ThomasA;Franceschi,Dinko;Dias,Mafalda;Frazer,Jonathan;Marks,DeboraS

文献摘要

相似文献

识别致病突变加速了遗传疾病的诊断和治疗开发。错义突变是基因诊断中的瓶颈,因为它们的影响不如截短或无义突变直接。虽然计算预测方法在预测已知疾病基因的变体方面越来越成功,但它们不能很好地推广到其他基因,因为分数没有在蛋白质组中校准1 -6。为了解决这个问题,我们开发了一个深度生成模型popEVE,它将进化信息与群体序列数据7相结合,并在按严重程度对变体进行排名方面实现了最先进的性能,以区分患有严重发育障碍的患者8和潜在的健康个体9。popEVE在这个发育障碍队列的患者中鉴定了442个基因,包括123种新型遗传性疾病的证据,其中许多不需要基因水平富集,也没有高估人群中致病性变异的患病率。这些变体中的大多数与3D复合物中的相互作用伙伴接近。对子外显子组的初步分析表明,popEVE可以识别候选变体,而不需要遗传标签。通过将变异放在一个统一的尺度上,我们的模型为整个蛋白质组和更广泛的人群中的适应性效应分布提供了一个全面的视角。popEVE为基因诊断提供了令人信服的证据,即使是在非常罕见的单一患者疾病中,依赖于重复观察的传统技术可能不适用。
Identifying causal mutations accelerates genetic disease diagnosis, and therapeutic development. Missense variants present a bottleneck in genetic diagnoses as their effects are less straightforward than truncations or nonsense mutations. While computational prediction methods are increasingly successful at prediction for variants in known disease genes, they do not generalize well to other genes as the scores are not calibrated across the proteome1–6. To address this, we developed a deep generative model, popEVE, that combines evolutionary information with population sequence data7 and achieves state-of-the-art performance at ranking variants by severity to distinguish patients with severe developmental disorders8 from potentially healthy individuals9. popEVE identifies 442 genes in patients this developmental disorder cohort, including evidence of 123 novel genetic disorders, many without the need for gene-level enrichment and without overestimating the prevalence of pathogenic variants in the population. A majority of these variants are close to interacting partners in 3D complexes. Preliminary analyses on child exomes indicate that popEVE can identify candidate variants without the need for inheritance labels. By placing variants on a unified scale, our model offers a comprehensive perspective on the distribution of fitness effects across the entire proteome and the broader human population. popEVE provides compelling evidence for genetic diagnoses even in exceptionally rare single-patient disorders where conventional techniques relying on repeated observations may not be applicable.