PROSET - A FAST PROCEDURE TO CREATE NONREDUNDANT SETS OF PROTEIN SEQUENCES

PROSET - A FAST PROCEDURE TO CREATE NONREDUNDANT SETS OF PROTEIN SEQUENCES
复制标题

DOI:
10.1016/0895-7177(92)90150-j
复制
发表时间:
1992-06-01
影响因子:
--
通讯作者:
BRENDEL, V
BRENDEL, V
中科院分区:
其他
文献类型:
--
作者:
BRENDEL, V

文献摘要

被引文献

相似文献

所描述的 PROSET 计算机程序有效地消除了一组蛋白质中的冗余条目。它发现原始组的所有蛋白质之间的重复,得出共享至少一个足够长的精确重复的任何两个序列之间的块同一性得分,并且如果得分超过用户定义的阈值,则丢弃两个序列中较短的一个。该程序可用于生成用于序列模式统计评估的控制集。它还可以用于减少可用数据库中的冗余量。
The PROSET computer program described efficiently eliminates redundant entries from a set of proteins. It finds repeats between all the proteins of the original set, derives a block identity score between any two sequences that share at least one sufficiently long exact repeat, and discards the shorter of the two sequences if the score exceeds a user-defined threshold. The program finds application in generating control sets for statistical evaluation of sequence patterns. It may also serve to reduce the amount of redundancy in available databases.