Mining for class-specific motifs in protein sequence classification.

Mining for class-specific motifs in protein sequence classification.
复制标题

DOI:
10.1186/1471-2105-14-96
复制
发表时间:
2013-03-15
期刊:
影响因子:
3
通讯作者:
Guda C
Guda C
中科院分区:
生物学4区
文献类型:
--
作者:
Srinivasan SM;Vural S;King BR;Guda C

文献摘要

参考文献

被引文献

相似文献

在蛋白质序列分类中,识别能够准确区分类别的序列基序或n元语法是一个比分类本身更有趣的科学问题。许多分类方法旨在进行准确的分类,但未能解释哪些序列特征确实有助于分类精度。我们假设,较低面额(n-gram)的序列可以用来探索序列图景,并在分类过程中识别区分类别的特定类别基序。区别性n-gram是在一个类别中出现频率很高,但在其他类别中存在或不存在的短肽序列。在这项研究中,我们提出了一种新的基于替换的评分函数,用于识别高度特定于一类的区别性n元语法。我们提出了一种基于区分n元语法的评分函数,可以有效地区分类别。评分功能最初从数据集中不同类别的蛋白质序列中获取整个4-8克的集合。将相同大小的相似n-gram组合成新的n-gram,其中相似性由BLOSUM62矩阵中的正氨基酸替代分数来定义。替代导致歧视性n-gram收获的数量大幅增加。由于数据集的不平衡性质,使用衰减因子来归一化n元语法的频率,这给予出现在较少类别中的n元语法更大的权重,反之亦然。在对n元语法进行归一化之后,计分函数为每个类别识别足够频繁以高于选择阈值的区别性4到8元语法。通过将这些有区别的n-gram映射回蛋白质序列,我们获得了连续的n-gram,它们代表了蛋白质序列中的短类别特定基序。与现有的主题发现方法Wordspy相比,我们的方法表现得很好。我们已经根据从NLSdb、ProSite和ELM数据库获得的重要功能基序验证了我们丰富的类特定基序集。我们证明了这种方法具有很强的通用性,因此可以广泛地应用于许多蛋白质序列分类任务中的特定类别基序的检测。所提出的评分函数和方法能够使用从蛋白质序列中获得的区别性n元语法来识别特定类别的基序。采用氨基酸替代分数进行相似性检测,采用抑制因子对不平衡数据集进行归一化处理,对评分函数的性能有显著影响。我们的多管齐下的验证测试表明,该方法可以从各种各样的蛋白质序列类别中检测到特定类别的基序,这对于检测不同生物的蛋白质组特定基序具有潜在的应用价值。
In protein sequence classification, identification of the sequence motifs or n-grams that can precisely discriminate between classes is a more interesting scientific question than the classification itself. A number of classification methods aim at accurate classification but fail to explain which sequence features indeed contribute to the accuracy. We hypothesize that sequences in lower denominations (n-grams) can be used to explore the sequence landscape and to identify class-specific motifs that discriminate between classes during classification. Discriminative n-grams are short peptide sequences that are highly frequent in one class but are either minimally present or absent in other classes. In this study, we present a new substitution-based scoring function for identifying discriminative n-grams that are highly specific to a class. We present a scoring function based on discriminative n-grams that can effectively discriminate between classes. The scoring function, initially, harvests the entire set of 4- to 8-grams from the protein sequences of different classes in the dataset. Similar n-grams of the same size are combined to form new n-grams, where the similarity is defined by positive amino acid substitution scores in the BLOSUM62 matrix. Substitution has resulted in a large increase in the number of discriminatory n-grams harvested. Due to the unbalanced nature of the dataset, the frequencies of the n-grams are normalized using a dampening factor, which gives more weightage to the n-grams that appear in fewer classes and vice-versa. After the n-grams are normalized, the scoring function identifies discriminative 4- to 8-grams for each class that are frequent enough to be above a selection threshold. By mapping these discriminative n-grams back to the protein sequences, we obtained contiguous n-grams that represent short class-specific motifs in protein sequences. Our method fared well compared to an existing motif finding method known as Wordspy. We have validated our enriched set of class-specific motifs against the functionally important motifs obtained from the NLSdb, Prosite and ELM databases. We demonstrate that this method is very generic; thus can be widely applied to detect class-specific motifs in many protein sequence classification tasks. The proposed scoring function and methodology is able to identify class-specific motifs using discriminative n-grams derived from the protein sequences. The implementation of amino acid substitution scores for similarity detection, and the dampening factor to normalize the unbalanced datasets have significant effect on the performance of the scoring function. Our multipronged validation tests demonstrate that this method can detect class-specific motifs from a wide variety of protein sequence classes with a potential application to detecting proteome-specific motifs of different organisms.
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.3233/spr-2008-0241
发表时间: 2008-01-01
影响因子: --
作者:
King, Brian R.;Guda, Chittibabu
通讯作者: Guda, Chittibabu
DOI: 10.1371/journal.pone.0027382
发表时间: 2011-11-10
期刊: PLOS ONE
影响因子: 3.7
作者:
Xiong, Hao;Capurso, Daniel;Segal, Mark R.
通讯作者: Segal, Mark R.
单词py:通过构建字典和学习语法来识别转录因子结合基序。
DOI: 10.1093/nar/gki492
发表时间: 2005-07-01
影响因子: 14.9
作者:
Wang G;Yu T;Zhang W
通讯作者: Zhang W
DOI: 10.1186/1471-2105-9-510
发表时间: 2008-12-01
期刊: BMC bioinformatics
影响因子: 3
作者:
Liu B;Wang X;Lin L;Dong Q;Wang X
通讯作者: Wang X