Mining sequence annotation databanks for association patterns

Mining sequence annotation databanks for association patterns
复制标题

DOI:
10.1093/bioinformatics/bti1206
复制
发表时间:
2005-01-01
期刊:
影响因子:
5.8
通讯作者:
Frishman, D
Frishman, D
中科院分区:
生物学3区
文献类型:
--
作者:
Artamonova, II;Frishman, G;Frishman, D

文献摘要

被引文献

相似文献

动机:目前将存放到序列数据库中的数百万个蛋白质序列将永远不会手动注释。自动软件管道产生的基于相似性的注释不可避免地包含了由于生物信息学方法的不完美而包含虚假分配。此类注释错误的示例包括使用固定识别阈值和错误的注释引起的过度预测和不正确的注释,这是由基于传播的信息传递到无关蛋白质或已在数据库中累积的错误转移的转移。生物信息学中最困难,最及时的挑战之一是智能系统的开发,旨在提高自动产生的注释的质量。解决此问题的一种可能的方法是基于关联规则挖掘的注释项中检测异常情况。分子:我们介绍了从两个大蛋白质注释数据库中得出的关联规则的第一个大规模分析,并揭示了小说,并揭示了小说,以前未知的规则强度分布趋势。大多数规则要么非常强大,要么非常虚弱,在中等强度范围内的规则相对较少。基于随后的瑞士 - 普罗特版本中的误差校正动力学,并且在我们自己的手动分析中,我们证明,强有力的规则的例外确实在注释误差中显着丰富,并且可以用来自动标记它们。我们确定了从瑞士 - 普罗特(Swiss-Prot)中不同领域得出的规则的不同强度依赖性。根据其组成项目而产生的关联规则的组成分解表明,可以纠正的大多数错误与基因功能角色有关。瑞士 - 普罗特的错误通常是由于其保守方法而导致的,而自动产生的送礼注释遭受了过量通道。
Motivation: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining.Results: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation.