Entropy-driven partitioning of the hierarchical protein space.

Entropy-driven partitioning of the hierarchical protein space.
复制标题

DOI:
10.1093/bioinformatics/btu478
复制
发表时间:
2014-09-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Linial M
Linial M
中科院分区:
其他
文献类型:
--
作者:
Rappoport N;Stern A;Linial N;Linial M

文献摘要

参考文献

被引文献

相似文献

Motivation: Modern protein sequencing techniques have led to the determination of >50 million protein sequences. ProtoNet is a clustering system that provides a continuous hierarchical agglomerative clustering tree for all proteins. While ProtoNet performs unsupervised classification of all included proteins, finding an optimal level of granularity for the purpose of focusing on protein functional groups remain elusive. Here, we ask whether knowledge-based annotations on protein families can support the automatic unsupervised methods for identifying high-quality protein families. We present a method that yields within the ProtoNet hierarchy an optimal partition of clusters, relative to manual annotation schemes. The method’s principle is to minimize the entropy-derived distance between annotation-based partitions and all available hierarchical partitions. We describe the best front (BF) partition of 2 478 328 proteins from UniRef50. Of 4 929 553 ProtoNet tree clusters, BF based on Pfam annotations contain 26 891 clusters. The high quality of the partition is validated by the close correspondence with the set of clusters that best describe thousands of keywords of Pfam. The BF is shown to be superior to naïve cut in the ProtoNet tree that yields a similar number of clusters. Finally, we used parameters intrinsic to the clustering process to enrich a priori the BF’s clusters. We present the entropy-based method’s benefit in overcoming the unavoidable limitations of nested clusters in ProtoNet. We suggest that this automatic information-based cluster selection can be useful for other large-scale annotation schemes, as well as for systematically testing and comparing putative families derived from alternative clustering methods. Availability and implementation: A catalog of BF clusters for thousands of Pfam keywords is provided at http://protonet.cs.huji.ac.il/bestFront/ Contact: michall@cc.huji.ac.il
DOI: 10.1093/nar/gks1211
发表时间: 2013-01
影响因子: 14.9
作者:
Sillitoe I;Cuff AL;Dessailly BH;Dawson NL;Furnham N;Lee D;Lees JG;Lewis TE;Studer RA;Rentzsch R;Yeats C;Thornton JM;Orengo CA
通讯作者: Orengo CA
DOI: 10.1093/nar/gkn762
发表时间: 2009-01
影响因子: 14.9
作者:
Wilson D;Pethica R;Zhou Y;Talbot C;Vogel C;Madera M;Chothia C;Gough J
通讯作者: Gough J
DOI: 10.1093/bioinformatics/bti542
发表时间: 2005-09-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Petryszak, R;Kretschmann, E;Apweiler, R
通讯作者: Apweiler, R
使用orthomcl将蛋白质分配给Orthomcl-DB组或将蛋白质组群集分为新的直系同源组。
DOI: 10.1002/0471250953.bi0612s35
发表时间: 2011-09
影响因子: --
作者:
Fischer, Steve;Brunk, Brian P;Chen, Feng;Gao, Xin;Harb, Omar S;Iodice, John B;Shanmugam, Dhanasekaran;Roos, David S;Stoeckert, Christian J Jr
通讯作者: Stoeckert, Christian J Jr
DOI: 10.1093/nar/29.1.29
发表时间: 2001-01-01
影响因子: 14.9
作者:
Barker, WC;Garavelli, JS;Wu, C
通讯作者: Wu, C