Identifying interaction sites in "recalcitrant" proteins: predicted protein and RNA binding sites in rev proteins of HIV-1 and EIAV agree with experimental data.

Identifying interaction sites in "recalcitrant" proteins: predicted protein and RNA binding sites in rev proteins of HIV-1 and EIAV agree with experimental data.
复制标题

DOI:
10.1142/9789812701626_0038
复制
发表时间:
2006-01-01
影响因子:
--
通讯作者:
Dobbs, Drena
Dobbs, Drena
中科院分区:
其他
文献类型:
--
作者:
Terribilini, Michael;Lee, Jae-Hyung;Dobbs, Drena

文献摘要

被引文献

相似文献

蛋白质-蛋白质和蛋白质核酸相互作用对于广泛的生物过程至关重要,包括基因表达的调节、蛋白质合成以及许多病毒的复制和组装。我们开发了机器学习方法,仅使用蛋白质序列作为输入来预测蛋白质的哪些氨基酸参与其与其他蛋白质和/或核酸的相互作用。在本文中,我们描述了在具有良好实验结构的蛋白质-蛋白质和蛋白质-RNA 复合物数据集上训练的分类器的应用。我们将这些分类器应用于预测临床重要蛋白质(其结构未知)序列中的蛋白质和 RNA 结合位点的问题:调节蛋白 Rev,对于 HIV-1 和其他慢病毒的复制至关重要。我们将我们的预测与已发表的 HIV-1 和 EIAV Rev 的生化、遗传和部分结构信息以及我们自己发表的 EIAV Rev 中 RNA 结合位点的实验图谱进行比较。预测的和实验确定的结合位点非常一致。可靠地预测直接促成特定结合事件的蛋白质残基的能力——无需有关蛋白质或其参与的复合物的结构信息——可能会产生新的疾病干预策略。
Protein-protein and protein nucleic acid interactions are vitally important for a wide range of biological processes, including regulation of gene expression, protein synthesis, and replication and assembly of many viruses. We have developed machine learning approaches for predicting which amino acids of a protein participate in its interactions with other proteins and/or nucleic acids, using only the protein sequence as input. In this paper, we describe an application of classifiers trained on datasets of well-characterized protein-protein and protein-RNA complexes for which experimental structures are available. We apply these classifiers to the problem of predicting protein and RNA binding sites in the sequence of a clinically important protein for which the structure is not known: the regulatory protein Rev, essential for the replication of HIV-1 and other lentiviruses. We compare our predictions with published biochemical, genetic and partial structural information for HIV-1 and EIAV Rev and with our own published experimental mapping of RNA binding sites in EIAV Rev. The predicted and experimentally determined binding sites are in very good agreement. The ability to predict reliably the residues of a protein that directly contribute to specific binding events--without the requirement for structural information regarding either the protein or complexes in which it participates--can potentially generate new disease intervention strategies.