IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders

IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders
复制标题

IMPROVE-DD:整合多种表型资源优化遗传决定的发育障碍的变异评估

DOI:
10.1101/2022.05.20.22275135
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Aitken S
Aitken S
中科院分区:
--
文献类型:
--
作者:
Aitken S

文献摘要

相似文献

使用全基因组测序数据诊断罕见发育障碍通常需要审查多个合理的候选变体,通常使用分类临床术语的本体。我们表明,整合多个表型资源优化变异评价发育障碍(IMPROVE-DD)通过纳入临床医生常用的数据和记录在健康记录中的其他类别。在这样做的过程中,我们量化的性别,生长和发展的独特贡献,除了人类表型本体论(HPO)的条款,并证明这些现成的信息来源的附加值。我们使用似然比的名义和定量数据,并提出了一个分类HPO条款在这个框架中。这种贝叶斯框架导致更稳健的诊断。使用在解密发育障碍研究中系统收集的数据,我们考虑了≥10个个体中具有致病性/可能致病性变体的77个基因。当对训练数据(AUC ≥ 0.6)进行测试时,所有基因都显示出至少令人满意的受试者操作特征预测,并且HPO项是大多数基因的最佳预测因子,尽管少数(13/77)基因被其他表型数据类型更好地预测。总体而言,基于多个综合表型数据源的分类器比基于任何单个来源的分类器表现更好,重要的是,综合模型产生的假阳性明显减少。最后,我们表明,具有良好的交叉验证预测性能的改进DD模型可以从相对较少的个人。这为候选基因的优先排序提出了新的策略,并强调了系统性临床数据收集对支持诊断程序的价值。
Diagnosing rare developmental disorders using genome-wide sequencing data commonly necessitates review of multiple plausible candidate variants, often using ontologies of categorical clinical terms. We show that Integrating Multiple Phenotype Resources Optimizes Variant Evaluation in Developmental Disorders (IMPROVE-DD) by incorporating additional classes of data commonly available to clinicians and recorded in health records. In doing so, we quantify the distinct contributions of sex, growth, and development in addition to Human Phenotype Ontology (HPO) terms and demonstrate added value from these readily available information sources. We use likelihood ratios for nominal and quantitative data and propose a classifier for HPO terms in this framework. This Bayesian framework results in more robust diagnoses. Using data systematically collected in the Deciphering Developmental Disorders study, we considered 77 genes with pathogenic/likely pathogenic variants in ≥10 individuals. All genes showed at least a satisfactory prediction by receiver operating characteristic when testing on training data (AUC ≥ 0.6), and HPO terms were the best predictor for the majority of genes, though a minority (13/77) of genes were better predicted by other phenotypic data types. Overall, classifiers based upon multiple integrated phenotypic data sources performed better than those based upon any individual source, and importantly, integrated models produced notably fewer false positives. Finally, we show that IMPROVE-DD models with good predictive performance on cross-validation can be constructed from relatively few individuals. This suggests new strategies for candidate gene prioritization and highlights the value of systematic clinical data collection to support diagnostic programs.