Identification of Divergent Protein Domains by Combining HMM-HMM Comparisons and Co-Occurrence Detection

Identification of Divergent Protein Domains by Combining HMM-HMM Comparisons and Co-Occurrence Detection
复制标题

DOI:
10.1371/journal.pone.0095275
复制
发表时间:
2014-06-05
期刊:
影响因子:
3.7
通讯作者:
Brehelin, Laurent
Brehelin, Laurent
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Ghouila, Amel;Florent, Isabelle;Brehelin, Laurent

文献摘要

被引文献

相似文献

蛋白质结构域的鉴定是了解蛋白质功能的关键步骤。隐马尔可夫模型(Hidden Markov Models,HMRM)已被证明是完成这一任务的有力工具。Pfam数据库特别提供了大量的被广泛用于测序生物体中蛋白质注释的HSPs集合。这是通过序列/HMM比较完成的。然而,这种方法可能缺乏敏感性时,搜索不同物种的域。最近,已经提出了HMM/HMM比较的方法,并证明在某些情况下比序列/HMM方法更灵敏。然而,这些方法通常不用于在基因组规模的蛋白质结构域发现,并没有调查,可以预期从他们的利用这个问题的好处。利用恶性疟原虫和L.主要的例子,我们调查在何种程度上HMM/HMM比较可以识别新的域出现还没有确定的序列/HMM的方法。我们发现,虽然HMM/HMM比较比序列/HMM比较敏感得多,但它们不足以准确地用作基因组规模的序列/HMM方法的独立补充。因此,我们建议使用域同现-一般域倾向于优先出现沿着与蛋白质中的一些喜爱的域-以提高该方法的准确性。我们表明,HMM/HMM比较和同现域检测的结合可以增强蛋白质注释。在5%的估计错误发现率下,它分别揭示了疟原虫和利什曼原虫蛋白中的901和1098个新结构域。人工检查这些预测的一部分表明,它包含了几个域的家庭,在这两个生物体中丢失。所有新的域出现已经被集成在EuPathDomains数据库中,沿着可以推导出的GO注释。
Identification of protein domains is a key step for understanding protein function. Hidden Markov Models (HMMs) have proved to be a powerful tool for this task. The Pfam database notably provides a large collection of HMMs which are widely used for the annotation of proteins in sequenced organisms. This is done via sequence/HMM comparisons. However, this approach may lack sensitivity when searching for domains in divergent species. Recently, methods for HMM/HMM comparisons have been proposed and proved to be more sensitive than sequence/HMM approaches in certain cases. However, these approaches are usually not used for protein domain discovery at a genome scale, and the benefit that could be expected from their utilization for this problem has not been investigated. Using proteins of P. falciparum and L. major as examples, we investigate the extent to which HMM/HMM comparisons can identify new domain occurrences not already identified by sequence/HMM approaches. We show that although HMM/HMM comparisons are much more sensitive than sequence/HMM comparisons, they are not sufficiently accurate to be used as a standalone complement of sequence/HMM approaches at the genome scale. Hence, we propose to use domain co-occurrence - the general domain tendency to preferentially appear along with some favorite domains in the proteins - to improve the accuracy of the approach. We show that the combination of HMM/HMM comparisons and co-occurrence domain detection boosts protein annotations. At an estimated False Discovery Rate of 5%, it revealed 901 and 1098 new domains in Plasmodium and Leishmania proteins, respectively. Manual inspection of part of these predictions shows that it contains several domain families that were missing in the two organisms. All new domain occurrences have been integrated in the EuPathDomains database, along with the GO annotations that can be deduced.