PredictSNP2: A Unified Platform for Accurately Evaluating SNP Effects by Exploiting the Different Characteristics of Variants in Distinct Genomic Regions.

PredictSNP2: A Unified Platform for Accurately Evaluating SNP Effects by Exploiting the Different Characteristics of Variants in Distinct Genomic Regions.
复制标题

DOI:
10.1371/journal.pcbi.1004962
复制
发表时间:
2016-05
影响因子:
4.3
通讯作者:
Brezovský J
Brezovský J
中科院分区:
生物学2区
文献类型:
--
作者:
Bendl J;Musil M;Štourač J;Zendulka J;Damborský J;Brezovský J

文献摘要

被引文献

相似文献

从人类基因组测序项目中获得的一个重要信息是,人类群体表现出约99.9%的遗传相似性。基因组其余部分的变异决定了我们的身份,追溯了我们的历史,揭示了我们的遗产。表型因果变异的精确描绘在提供遗传性疾病的准确个性化诊断、预后和治疗中起着关键作用。最近已经报道了几种实现这种划分的计算方法。然而,他们的能力,以查明潜在的有害变异是有限的事实,他们的预测机制不占不同类别的变异的存在。因此,它们的输出偏向于在变异数据库中表现最强烈的变异类别。此外,大多数这样的方法提供数值分数,但不提供对变体的危险性的二进制预测或更容易被用户理解的置信度分数。我们构建了三个数据集,涵盖不同类型的疾病相关变体,分为五类:(i)调节,(ii)剪接,(iii)错义,(iv)同义和(v)无义变体。这些数据集用于开发类别最佳决策阈值,并评估用于变体优先级排序的六种工具:CADD,DANN,FATHMM,FitCons,FunSeq2和GWAVA。这一评价揭示了基于类别的方法的一些重要优点。然后将使用五种性能最佳的工具获得的结果合并为共识评分。额外的比较分析表明,在错义变异的情况下,基于蛋白质的预测因子比基于DNA序列的预测因子表现得更好。开发了一个用户友好的网络界面,可以方便地访问五个工具的预测及其共识得分,其格式是针对不同类别变化的具体特征而定制的,用户可以理解。为了对变体进行全面评估,预测还补充了来自八个数据库的注释。该网络服务器可在www.example.com上免费提供给社区。
An important message taken from human genome sequencing projects is that the human population exhibits approximately 99.9% genetic similarity. Variations in the remaining parts of the genome determine our identity, trace our history and reveal our heritage. The precise delineation of phenotypically causal variants plays a key role in providing accurate personalized diagnosis, prognosis, and treatment of inherited diseases. Several computational methods for achieving such delineation have been reported recently. However, their ability to pinpoint potentially deleterious variants is limited by the fact that their mechanisms of prediction do not account for the existence of different categories of variants. Consequently, their output is biased towards the variant categories that are most strongly represented in the variant databases. Moreover, most such methods provide numeric scores but not binary predictions of the deleteriousness of variants or confidence scores that would be more easily understood by users. We have constructed three datasets covering different types of disease-related variants, which were divided across five categories: (i) regulatory, (ii) splicing, (iii) missense, (iv) synonymous, and (v) nonsense variants. These datasets were used to develop category-optimal decision thresholds and to evaluate six tools for variant prioritization: CADD, DANN, FATHMM, FitCons, FunSeq2 and GWAVA. This evaluation revealed some important advantages of the category-based approach. The results obtained with the five best-performing tools were then combined into a consensus score. Additional comparative analyses showed that in the case of missense variations, protein-based predictors perform better than DNA sequence-based predictors. A user-friendly web interface was developed that provides easy access to the five tools’ predictions, and their consensus scores, in a user-understandable format tailored to the specific features of different categories of variations. To enable comprehensive evaluation of variants, the predictions are complemented with annotations from eight databases. The web server is freely available to the community at http://loschmidt.chemi.muni.cz/predictsnp2.