NrichD database: sequence databases enriched with computationally designed protein-like sequences aid in remote homology detection.

NrichD database: sequence databases enriched with computationally designed protein-like sequences aid in remote homology detection.
复制标题

DOI:
10.1093/nar/gku888
复制
发表时间:
2015-01
影响因子:
14.9
通讯作者:
Srinivasan N
Srinivasan N
中科院分区:
生物学2区
文献类型:
--
作者:
Mudgal R;Sandhya S;Kumar G;Sowdhamini R;Chandra NR;Srinivasan N

文献摘要

参考文献

被引文献

相似文献

NrichD(http://proline.biochem.iisc.ernet.in/NRICHD/)是一个计算设计的蛋白质样序列的数据库,被扩充到天然序列数据库中,可以在蛋白质序列空间中进行跳跃,以帮助检测远程关系。在缺乏结构证据或天然“中间相关序列”的情况下建立蛋白质关系是一项具有挑战性的任务。最近,我们已经证明,人工中间序列/接头的计算设计是一种有效的方法来填补蛋白质序列空间中天然存在的空隙。通过大规模的评估,我们已经证明,这些序列可以插入到常用的搜索数据库,以提高性能的常规使用的序列搜索方法在检测远程关系。由于预期这些数据集将用于建立蛋白质关系,已经在结构和功能域水平捕获这些关系的两个数据库,即SCOP数据库和Pfam数据库,已经用这些人工中间序列“富集”。NrichD数据库目前包含3611010个人工序列,这些序列来自374个SCOP折叠,在27882对家族之间产生。数据集可免费下载。其他功能包括设计用户感兴趣的任何两个蛋白质家族之间的人工序列。
NrichD (http://proline.biochem.iisc.ernet.in/NRICHD/) is a database of computationally designed protein-like sequences, augmented into natural sequence databases that can perform hops in protein sequence space to assist in the detection of remote relationships. Establishing protein relationships in the absence of structural evidence or natural ‘intermediately related sequences’ is a challenging task. Recently, we have demonstrated that the computational design of artificial intermediary sequences/linkers is an effective approach to fill naturally occurring voids in protein sequence space. Through a large-scale assessment we have demonstrated that such sequences can be plugged into commonly employed search databases to improve the performance of routinely used sequence search methods in detecting remote relationships. Since it is anticipated that such data sets will be employed to establish protein relationships, two databases that have already captured these relationships at the structural and functional domain level, namely, the SCOP database and the Pfam database, have been ‘enriched’ with these artificial intermediary sequences. NrichD database currently contains 3 611 010 artificial sequences that have been generated between 27 882 pairs of families from 374 SCOP folds. The data sets are freely available for download. Additional features include the design of artificial sequences between any two protein families of interest to the user.
DOI: 10.1006/jmbi.1999.2653
发表时间: 1999-04-16
影响因子: 5.6
作者:
Aravind, L;Koonin, EV
通讯作者: Koonin, EV
DOI: 10.1093/bioinformatics/btm034
发表时间: 2007-04-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bateman, Alex;Finn, Robert D.
通讯作者: Finn, Robert D.
DOI: 10.1126/science.278.5335.82
发表时间: 1997-10-03
期刊: SCIENCE
影响因子: 56.9
作者:
Dahiyat, BI;Mayo, SL
通讯作者: Mayo, SL
DOI: 10.1038/nature03991
发表时间: 2005-09-22
期刊: NATURE
影响因子: 64.8
作者:
Socolich, M;Lockless, SW;Ranganathan, R
通讯作者: Ranganathan, R
DOI: 10.1093/protein/12.2.95
发表时间: 1999-02-01
期刊: PROTEIN ENGINEERING
影响因子: --
作者:
Salamov, AA;Suwa, M;Swindells, MB
通讯作者: Swindells, MB