Predicting sequence and structural specificities of RNA binding regions recognized by splicing factor SRSF1.

Predicting sequence and structural specificities of RNA binding regions recognized by splicing factor SRSF1.
复制标题

DOI:
10.1186/1471-2164-12-s5-s8
复制
发表时间:
2011-12-23
期刊:
影响因子:
4.4
通讯作者:
Liu Y
Liu Y
中科院分区:
生物学2区
文献类型:
--
作者:
Wang X;Juan L;Lv J;Wang K;Sanford JR;Liu Y

文献摘要

被引文献

相似文献

RNA结合蛋白在真核生物RNA加工过程中发挥着重要作用。尽管它们在编码和非编码RNA生物发生和调控中具有普遍的功能,但阐明定义蛋白质-RNA相互作用的序列特异性仍然是一个重大挑战。最近,CLIP-seq(交联免疫沉淀,随后进行高通量测序)已成功实施,以研究SRSF 1,PTBP 1,NOVA和fox 2蛋白的转录组范围的结合模式。这些研究要么采用传统的方法,如多重EM的模体激发(MEME),以发现序列的一致性RBP的结合位点或使用Z-评分统计,以寻找一定大小的过度代表的核苷酸。我们认为,大多数这些方法都不适合RNA基序鉴定,因为它们无法纳入蛋白质-RNA相互作用的RNA结构背景,这可能会影响结合特异性。在这里,我们描述了一种新的基于模型的方法-RNAMotifModeler,通过整合序列特征和RNA二级结构来识别蛋白质-RNA结合区域的一致性。作为一个例子,我们在SRSF 1(SF 2/ASF)CLIP-seq数据上实现了RNAMotifModeler。我们鉴定的序列结构一致性是在高度单链RNA背景下的富含嘌呤的八聚体“AGAAGAAG”。未配对的概率,即不形成配对的概率,显著高于阴性对照和结合位点周围的侧翼序列,表明SRSF 1蛋白倾向于结合单链RNA。进一步的统计分析表明,SRSF 1 octamer motif的第二和第五碱基具有更强的序列特异性,但单链性较弱;第三、第四、第六和第七碱基更可能是单链的,但具有更强的简并序列特异性。因此,我们推测,核苷酸特异性和二级结构发挥互补作用,在结合位点识别SRSF 1。在这项研究中,我们提出了一个计算模型来预测蛋白质-RNA结合区域的序列一致性和最佳RNA二级结构。SRSF 1 CLIP-seq数据的成功实施证明了提高我们对RNA结合蛋白结合特异性的理解的巨大潜力。
RNA-binding proteins (RBPs) play diverse roles in eukaryotic RNA processing. Despite their pervasive functions in coding and noncoding RNA biogenesis and regulation, elucidating the sequence specificities that define protein-RNA interactions remains a major challenge. Recently, CLIP-seq (Cross-linking immunoprecipitation followed by high-throughput sequencing) has been successfully implemented to study the transcriptome-wide binding patterns of SRSF1, PTBP1, NOVA and fox2 proteins. These studies either adopted traditional methods like Multiple EM for Motif Elicitation (MEME) to discover the sequence consensus of RBP's binding sites or used Z-score statistics to search for the overrepresented nucleotides of a certain size. We argue that most of these methods are not well-suited for RNA motif identification, as they are unable to incorporate the RNA structural context of protein-RNA interactions, which may affect to binding specificity. Here, we describe a novel model-based approach--RNAMotifModeler to identify the consensus of protein-RNA binding regions by integrating sequence features and RNA secondary structures. As an example, we implemented RNAMotifModeler on SRSF1 (SF2/ASF) CLIP-seq data. The sequence-structural consensus we identified is a purine-rich octamer 'AGAAGAAG' in a highly single-stranded RNA context. The unpaired probabilities, the probabilities of not forming pairs, are significantly higher than negative controls and the flanking sequence surrounding the binding site, indicating that SRSF1 proteins tend to bind on single-stranded RNA. Further statistical evaluations revealed that the second and fifth bases of SRSF1octamer motif have much stronger sequence specificities, but weaker single-strandedness, while the third, fourth, sixth and seventh bases are far more likely to be single-stranded, but have more degenerate sequence specificities. Therefore, we hypothesize that nucleotide specificity and secondary structure play complementary roles during binding site recognition by SRSF1. In this study, we presented a computational model to predict the sequence consensus and optimal RNA secondary structure for protein-RNA binding regions. The successful implementation on SRSF1 CLIP-seq data demonstrates great potential to improve our understanding on the binding specificity of RNA binding proteins.