Affinity regression predicts the recognition code of nucleic acid-binding proteins.

Affinity regression predicts the recognition code of nucleic acid-binding proteins.
复制标题

DOI:
10.1038/nbt.3343
复制
发表时间:
2015-12
影响因子:
46.9
通讯作者:
Leslie CS
Leslie CS
中科院分区:
工程技术1区
文献类型:
--
作者:
Pelossof R;Singh I;Yang JL;Weirauch MT;Hughes TR;Leslie CS

文献摘要

被引文献

相似文献

直接从蛋白质序列预测核酸结合蛋白的亲和力是一个主要的未解决的问题。我们提出了一种统计学方法,用于从高通量结合分析中学习转录因子(TF)或RNA结合蛋白(RBP)家族的识别代码。我们的方法,称为亲和力回归,在蛋白质结合微阵列(PBM)或RNA竞争实验上进行训练,以学习蛋白质和核酸之间的相互作用模型,仅使用蛋白质结构域和探针序列作为输入。通过对小鼠同源结构域PBM谱的训练,我们的模型正确地识别了赋予DNA结合特异性的残基,并准确地预测了一组独立的不同同源结构域的结合基序。类似地,从RNA竞争配置文件学习不同的RBP,我们的模型可以预测保持出来的蛋白质的结合亲和力,并确定关键的RNA结合残基。更广泛地说,我们设想应用我们的方法来建模和预测生物相互作用的任何设置,其中有一个高通量的“亲和力”读出。
Predicting the affinity profiles of nucleic acid-binding proteins directly from the protein sequence is a major unsolved problem. We present a statistical approach for learning the recognition code of a family of transcription factors (TFs) or RNA-binding proteins (RBPs) from high-throughput binding assays. Our method, called affinity regression, trains on protein binding microarray (PBM) or RNA compete experiments to learn an interaction model between proteins and nucleic acids, using only protein domain and probe sequences as inputs. By training on mouse homeodomain PBM profiles, our model correctly identifies residues that confer DNA-binding specificity and accurately predicts binding motifs for an independent set of divergent homeodomains. Similarly, learning from RNA compete profiles for diverse RBPs, our model can predict the binding affinities of held-out proteins and identify key RNA-binding residues. More broadly, we envision applying our method to model and predict biological interactions in any setting where there is a high-throughput ‘affinity’ readout.