eBLOCKs: enumerating conserved protein blocks to achieve maximal sensitivity and specificity.

eBLOCKs: enumerating conserved protein blocks to achieve maximal sensitivity and specificity.
复制标题

DOI:
10.1093/nar/gki060
复制
发表时间:
2005-01-01
影响因子:
14.9
通讯作者:
Brutlag DL
Brutlag DL
中科院分区:
生物学2区
文献类型:
--
作者:
Su QJ;Lu L;Saxonov S;Brutlag DL

文献摘要

参考文献

被引文献

相似文献

将蛋白质分为家族和超家族可以识别具有重要功能的保守结构域。来自这些保守区域的基序和评分矩阵提供了识别新序列中相似模式的计算工具,从而能够预测基因组的蛋白质功能。EBLOCKs数据库列举了一系列蛋白质块,每个功能域具有不同的保守水平。一个生物学上重要的区域在高度相似的较小的蛋白质家族中是最严格保守的。相同的区域通常在一组更远的相关蛋白质中被发现,但严密性降低。通过枚举,可以从列更多、家庭成员更少的块生成高度特定的签名,而高度敏感的签名可以从列更少、成员更多的块中派生出来,就像在超家族中一样。通过应用PSI-BLAST和改进的K-Means聚类算法,eBLOCKs根据不同的相似度自动对蛋白质序列进行分组。进行多个序列比对,并将其修剪成一系列未映射的块。基序和特定位置的评分矩阵来自eBLOCK,并可用于序列搜索和注释。EBLOCKs数据库为高通量基因组注释提供了一个工具,具有最大的特异性和敏感度。EBLOCKS数据库可在万维网上免费获得,网址为http://motif.stanford.edu/eblocks/,供所有用户在线使用。希望获得该计划副本的学术和非营利机构可联系道格拉斯·L·布鲁特拉格(Brutlag@stanford.edu)。希望在内部安装该程序副本的商业公司可联系斯坦福大学技术许可办公室的Jacqueline Tay(jacqueline.tay@stanford.edu;http://otl.stanford.edu/).
Classifying proteins into families and superfamilies allows identification of functionally important conserved domains. The motifs and scoring matrices derived from such conserved regions provide computational tools that recognize similar patterns in novel sequences, and thus enable the prediction of protein function for genomes. The eBLOCKs database enumerates a cascade of protein blocks with varied conservation levels for each functional domain. A biologically important region is most stringently conserved among a smaller family of highly similar proteins. The same region is often found in a larger group of more remotely related proteins with a reduced stringency. Through enumeration, highly specific signatures can be generated from blocks with more columns and fewer family members, while highly sensitive signatures can be derived from blocks with fewer columns and more members as in a superfamily. By applying PSI-BLAST and a modified K-means clustering algorithm, eBLOCKs automatically groups protein sequences according to different levels of similarity. Multiple sequence alignments are made and trimmed into a series of ungapped blocks. Motifs and position-specific scoring matrices were derived from eBLOCKs and made available for sequence search and annotation. The eBLOCKs database provides a tool for high-throughput genome annotation with maximal specificity and sensitivity. The eBLOCKs database is freely available on the World Wide Web at http://motif.stanford.edu/eblocks/ to all users for online usage. Academic and not-for-profit institutions wishing copies of the program may contact Douglas L. Brutlag (brutlag@stanford.edu). Commercial firms wishing copies of the program for internal installation may contact Jacqueline Tay at the Stanford Office of Technology Licensing (jacqueline.tay@stanford.edu; http://otl.stanford.edu/).
DOI: 10.1093/nar/28.1.270
发表时间: 2000-01-01
影响因子: 14.9
作者:
Krause, A;Stoye, J;Vingron, M
通讯作者: Vingron, M
DOI: 10.1093/nar/gkp985
发表时间: 2010-01
影响因子: 14.9
作者:
Finn RD;Mistry J;Tate J;Coggill P;Heger A;Pollington JE;Gavin OL;Gunasekaran P;Ceric G;Forslund K;Holm L;Sonnhammer EL;Eddy SR;Bateman A
通讯作者: Bateman A
DOI: 10.1093/nar/gkh039
发表时间: 2004-01-01
影响因子: 14.9
作者:
Andreeva, A;Howorth, D;Murzin, AG
通讯作者: Murzin, AG
DOI: 10.1093/nar/28.1.49
发表时间: 2000-01-01
影响因子: 14.9
作者:
Yona, G;Linial, N;Linial, M
通讯作者: Linial, M
DOI: 10.1093/nar/gkg035
发表时间: 2003-01-01
影响因子: 14.9
作者:
Kriventseva, EV;Servant, F;Apweiler, R
通讯作者: Apweiler, R