The use of gene ontology evidence codes in preventing classifier assessment bias

The use of gene ontology evidence codes in preventing classifier assessment bias
复制标题

DOI:
10.1093/bioinformatics/btp122
复制
发表时间:
2009-05-01
期刊:
影响因子:
5.8
通讯作者:
Ben-Hur, Asa
Ben-Hur, Asa
中科院分区:
生物学3区
文献类型:
--
作者:
Rogers, Mark F.;Ben-Hur, Asa

文献摘要

被引文献

相似文献

动机:生物界对蛋白质功能的计算注释的依赖使得功能预测方法的正确评估成为一个非常重要的问题。当前生物数据库中的大部分注释是基于计算方法的,这一事实可能导致在估计功能预测方法的准确性时存在偏差。这可能会发生,因为预测的注释,计算得出的摆在首位可能比预测注释,实验得出的,导致过于乐观的分类器性能estimates.Results:我们说明了这种现象在一组控制实验中使用最近邻分类器,使用PSI-BLAST相似性分数。我们的研究结果表明,用于评估蛋白质功能预测因子的基因本体(GO)注释的来源可能对分类器的准确性产生非常显著的影响:当分类器被给予对注释的访问权时,生物过程命名空间中的四个物种和GO项的平均准确度从0.72增加到0.87,所述注释被分配了指示可能的计算源的证据代码,而不是实验确定的注释。在其他命名空间中观察到的增长幅度稍小。在这些比较中的注释和他们的分布在GO terms的总数保持不变。结论:总之,考虑到GO证据代码是需要报告的准确性统计数据,不高估模型的性能,是一个公平的比较,依赖于不同的信息来源的分类特别重要。
Motivation: The biological community's reliance on computational annotations of protein function makes correct assessment of function prediction methods an issue of great importance. The fact that a large fraction of the annotations in current biological databases are based on computational methods can lead to bias in estimating the accuracy of function prediction methods. This can happen since predicting an annotation that was derived computationally in the first place is likely easier than predicting annotations that were derived experimentally, leading to over-optimistic classifier performance estimates.Results: We illustrate this phenomenon in a set of controlled experiments using a nearest neighbor classifier that uses PSI-BLAST similarity scores. Our results demonstrate that the source of Gene Ontology (GO) annotations used to assess a protein function predictor can have a highly significant influence on classifier accuracy: the average accuracy over four species and over GO terms in the biological process namespace increased from 0.72 to 0.87 when the classifier was given access to annotations that are assigned evidence codes that indicate a possible computational source, instead of experimentally determined annotations. Slightly smaller increases were observed in the other namespaces. In these comparisons the total number of annotations and their distribution across GO terms were kept the same.Conclusion: In conclusion, taking into account GO evidence codes is required for reporting accuracy statistics that do not overestimate a model's performance, and is of particular importance for a fair comparison of classifiers that rely on different information sources.