Automatic rule generation for protein annotation with the C4.5 data mining algorithm applied on SWISS-PROT

Automatic rule generation for protein annotation with the C4.5 data mining algorithm applied on SWISS-PROT
复制标题

DOI:
10.1093/bioinformatics/17.10.920
复制
发表时间:
2001-10-01
期刊:
影响因子:
5.8
通讯作者:
Apweiler, R
Apweiler, R
中科院分区:
生物学3区
文献类型:
--
作者:
Kretschmann, E;Fleischmann, W;Apweiler, R

文献摘要

被引文献

相似文献

动机:新提交的蛋白质数据量与公共数据库中可靠的功能注释之间的差距越来越大。传统的人工注释文献,策展和序列分析工具,而不使用自动注释系统是无法跟上不断增加的数据量提交。对人工管理的数据库(如TrEMBL或GenPept)的自动补充涵盖了原始数据,但仅提供了有限的注释。为了改善这种情况,需要自动工具,支持手动注释,自动增加可靠的信息量,并帮助检测手动生成的annotations.Results不一致:一个标准的数据挖掘算法成功地应用于获得知识的关键字标注SWISS-PROT。生成了11306条规则,这些规则在数据库中提供,并且可以应用于尚未注释的蛋白质序列并使用网络浏览器查看。它们依赖于发现蛋白质的生物体的分类学及其序列的签名匹配。通过交叉验证生成的规则的统计评估表明,通过将它们应用于任意蛋白质,可以生成33%的关键字注释,错误率为1.5%。通过容忍5%的更高错误率,关键字注释的覆盖率可以增加到60%。
Motivation: The gap between the amount of newly submitted protein data and reliable functional annotation in public databases is growing. Traditional manual annotation by literature, curation and sequence analysis tools without the use of automated annotation systems is not able to keep up with the ever increasing quantity of data that is submitted. Automated supplements to manually curated databases such as TrEMBL or GenPept cover raw data but provide only limited annotation. To improve this situation automatic tools are needed that support manual annotation, automatically increase the amount of reliable information and help to detect inconsistencies in manually generated annotations.Results: A standard data mining algorithm was successfully applied to gain knowledge about the Keyword annotation in SWISS-PROT. 11306 rules were generated, which are, provided in a database and can be applied to yet unannotated protein sequences and viewed using a web browser. They rely on the taxonomy of the organism, in which the protein was found and on signature matches of its, sequence. The statistical evaluation of the generated rules by cross-validation suggests that by applying them on arbitrary proteins 33% of their keyword annotation can be generated with an error rate of 1.5%. The coverage rate of the keyword annotation can be, increased to 60% by tolerating a higher error rate of 5%.