A comparative study of machine-learning methods to predict the effects of single nucleotide polymorphisms on protein function

A comparative study of machine-learning methods to predict the effects of single nucleotide polymorphisms on protein function
复制标题

DOI:
10.1093/bioinformatics/btg297
复制
发表时间:
2003-11-22
期刊:
影响因子:
5.8
通讯作者:
Westhead, DR
Westhead, DR
中科院分区:
生物学3区
文献类型:
--
作者:
Krishnan, VG;Westhead, DR

文献摘要

被引文献

相似文献

动机:现在可用的大量单核苷酸多态数据促使方法的发展,以区分中性变化和那些具有真正生物影响的变化。在这里,两种不同的机器学习方法--决策树和支持向量机首次被应用于这个问题。与大多数其他方法一样,只考虑了基因组蛋白质编码区的非同义变化。结果:在详细的交叉验证分析中,两种学习方法都被证明与现有方法竞争良好,并在一些关键测试中表现优于它们。支持向量机表现出更好的泛化性能,但决策树的优势是生成具有预测置信度稳健估计的可解释规则。结果表明,蛋白质结构信息的包含产生了更准确的方法,这与最近的其他研究一致,并评估了使用预测结构而不是实际结构的效果。
Motivation: The large volume of single nucleotide polymorphism data now available motivates the development of methods for distinguishing neutral changes from those which have real biological effects. Here, two different machine-learning methods, decision trees and support vector machines (SVMs), are applied for the first time to this problem. In common with most other methods, only non-synonymous changes in protein coding regions of the genome are considered.Results: In detailed cross-validation analysis, both learning methods are shown to compete well with existing methods, and to out-perform them in some key tests. SVMs show better generalization performance, but decision trees have the advantage of generating interpretable rules with robust estimates of prediction confidence. It is shown that the inclusion of protein structure information produces more accurate methods, in agreement with other recent studies, and the effect of using predicted rather than actual structure is evaluated.