Predicting substrate specificity of adenylation domains of nonribosomal peptide synthetases and other protein properties by latent semantic indexing

Predicting substrate specificity of adenylation domains of nonribosomal peptide synthetases and other protein properties by latent semantic indexing
复制标题

DOI:
10.1007/s10295-013-1322-2
复制
发表时间:
2014-02-01
影响因子:
3.4
通讯作者:
Starcevic, Antonio
Starcevic, Antonio
中科院分区:
工程技术3区
文献类型:
--
作者:
Baranasic, Damir;Zucko, Jurica;Starcevic, Antonio

文献摘要

被引文献

相似文献

成功的基因组挖掘依赖于从序列中准确预测蛋白质功能。这通常涉及将蛋白质家族分成功能亚型(例如,具有不同的衬底)。在许多情况下,只有少数已知的功能亚型,但在非核糖体肽合成酶(NRPS)的腺苷酸化结构域的情况下,有超过500种已知的底物。潜在语义索引(LSI)最初是为文本处理而开发的,但也用于将蛋白质分配到家族。蛋白质被视为“文档”,并且有必要将氨基酸序列的属性编码为“术语”,以构建术语-文档矩阵,该矩阵对每个文档中的术语进行计数。然后处理该矩阵以产生文档概念矩阵,其中每个蛋白质表示为行向量。向量彼此接近程度的标准度量(向量之间夹角的余弦)提供了蛋白质相似性的度量。以前的工作编码蛋白质作为寡肽术语,即计数寡肽,但没有使用关于寡肽在蛋白质中的位置的信息。提出了一种新的标记化方法来分析多重比对信息。LSI成功地区分了两个功能亚型,在五个良好的特点家庭。不同“概念”维度的可视化允许探索蛋白质家族的结构。LSI也被用来预测NRPS的腺苷酸化结构域的氨基酸底物。当使用来自多重比对的选定残基而不是腺苷酸化结构域的总序列时,获得了更好的结果。使用来自底物结合口袋的10个残基比使用8个活性位点内的34个残基表现更好。预测效率略优于最好的出版方法,使用支持向量机。
Successful genome mining is dependent on accurate prediction of protein function from sequence. This often involves dividing protein families into functional subtypes (e.g., with different substrates). In many cases, there are only a small number of known functional subtypes, but in the case of the adenylation domains of nonribosomal peptide synthetases (NRPS), there are > 500 known substrates. Latent semantic indexing (LSI) was originally developed for text processing but has also been used to assign proteins to families. Proteins are treated as ''documents'' and it is necessary to encode properties of the amino acid sequence as ''terms'' in order to construct a term-document matrix, which counts the terms in each document. This matrix is then processed to produce a document-concept matrix, where each protein is represented as a row vector. A standard measure of the closeness of vectors to each other (cosines of the angle between them) provides a measure of protein similarity. Previous work encoded proteins as oligopeptide terms, i.e. counted oligopeptides, but used no information regarding location of oligopeptides in the proteins. A novel tokenization method was developed to analyze information from multiple alignments. LSI successfully distinguished between two functional subtypes in five well-characterized families. Visualization of different ''concept'' dimensions allows exploration of the structure of protein families. LSI was also used to predict the amino acid substrate of adenylation domains of NRPS. Better results were obtained when selected residues from multiple alignments were used rather than the total sequence of the adenylation domains. Using ten residues from the substrate binding pocket performed better than using 34 residues within 8 of the active site. Prediction efficiency was somewhat better than that of the best published method using a support vector machine.