CADD v1.7: using protein language models, regulatory CNNs and other nucleotide-level scores to improve genome-wide variant predictions.

CADD v1.7: using protein language models, regulatory CNNs and other nucleotide-level scores to improve genome-wide variant predictions.
复制标题

CADD v1.7:使用蛋白质语言模型、调控CNN和其他核苷酸水平评分来改善全基因组变异预测。

DOI:
10.1093/nar/gkad989
复制
发表时间:
2024-01-05
影响因子:
14.9
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

相似文献

基于机器学习的遗传变异评分和分类有助于临床结果的评估,并用于在各种遗传研究和分析中确定变异的优先级。联合注释依赖缺失(Combined Annotation-Dependent Depletion, CADD)是最早用于跨不同分子功能的变异全基因组优先排序的方法之一,自最初发表以来一直在不断发展和完善。在这里,我们介绍我们的最新版本,CADD v1.7。我们探索并整合了新的注释功能,其中包括最先进的蛋白质语言模型评分(Meta ESM-1v),调节变异效应预测(来自基于序列的卷积神经网络)和序列保守评分(Zoonomia)。我们在ClinVar、ExAC/gnomAD和1000个基因组变体的数据集上评估了新版本。对于编码效应,我们在ProteinGym的31个深度突变扫描(DMS)数据集上测试了CADD,对于调控效应预测,我们使用了启动子和增强子序列的饱和突变报告分析数据。新特性的加入进一步提高了CADD的整体性能。与之前的版本一样,所有数据集、全基因组CADD v1.7分数、现场评分脚本和易于使用的web服务器都可以通过https://cadd.bihealth.org/或https://cadd.gs.washington.edu/随时提供给社区。
Machine Learning-based scoring and classification of genetic variants aids the assessment of clinical findings and is employed to prioritize variants in diverse genetic studies and analyses. Combined Annotation-Dependent Depletion (CADD) is one of the first methods for the genome-wide prioritization of variants across different molecular functions and has been continuously developed and improved since its original publication. Here, we present our most recent release, CADD v1.7. We explored and integrated new annotation features, among them state-of-the-art protein language model scores (Meta ESM-1v), regulatory variant effect predictions (from sequence-based convolutional neural networks) and sequence conservation scores (Zoonomia). We evaluated the new version on data sets derived from ClinVar, ExAC/gnomAD and 1000 Genomes variants. For coding effects, we tested CADD on 31 Deep Mutational Scanning (DMS) data sets from ProteinGym and, for regulatory effect prediction, we used saturation mutagenesis reporter assay data of promoter and enhancer sequences. The inclusion of new features further improved the overall performance of CADD. As with previous releases, all data sets, genome-wide CADD v1.7 scores, scripts for on-site scoring and an easy-to-use webserver are readily provided via https://cadd.bihealth.org/ or https://cadd.gs.washington.edu/ to the community.