Redefining the structural motifs that determine RNA binding and RNA editing by pentatricopeptide repeat proteins in land plants

Redefining the structural motifs that determine RNA binding and RNA editing by pentatricopeptide repeat proteins in land plants
复制标题

DOI:
10.1111/tpj.13121
复制
发表时间:
2016-02-01
期刊:
影响因子:
7.2
通讯作者:
Small, Ian
Small, Ian
中科院分区:
生物学1区
文献类型:
--
作者:
Cheng, Shifeng;Gutmann, Bernard;Small, Ian

文献摘要

被引文献

相似文献

五肽重复序列(PPR)是陆地植物中最大的蛋白质家族之一。它们的特征是串联30-40个氨基酸基序,形成一个延伸的结合表面,能够识别RNA链的序列特异性。它们几乎都是翻译后的靶标,在转录后过程中发挥着重要的作用,包括剪接、RNA编辑和翻译的启动。描述PPR蛋白质如何识别其RNA靶标的代码有望加快对这些蛋白质的研究,但利用这种代码需要对每种蛋白质中所有不同的核苷酸结合基序进行准确的定义和注释。我们使用结构建模的方法定义了在植物蛋白中发现的PPR基序的10种不同变体,以及在许多RNA编辑因子的C末端发现的假定的脱氨酶基序。我们表明,RNA编辑因子的超螺旋RNA结合面可能比以前认识的更长。我们使用重新定义的基序对109个基因组中的PPR序列进行了准确和一致的注释。我们报道,由于基因融合和虚假内含子的插入,在许多公共植物蛋白质组中,PPR基因模型的错误率很高。这些广泛物种的一致注释的数据集是未来比较基因组学研究的宝贵资源,也是准确的大规模PPR目标计算预测的必要先决条件。我们创建了一个门户网站(),为社区提供对这些资源的开放访问。
The pentatricopeptide repeat (PPR) proteins form one of the largest protein families in land plants. They are characterised by tandem 30-40 amino acid motifs that form an extended binding surface capable of sequence-specific recognition of RNA strands. Almost all of them are post-translationally targeted to plastids and mitochondria, where they play important roles in post-transcriptional processes including splicing, RNA editing and the initiation of translation. A code describing how PPR proteins recognise their RNA targets promises to accelerate research on these proteins, but making use of this code requires accurate definition and annotation of all of the various nucleotide-binding motifs in each protein. We have used a structural modelling approach to define 10 different variants of the PPR motif found in plant proteins, in addition to the putative deaminase motif that is found at the C-terminus of many RNA-editing factors. We show that the super-helical RNA-binding surface of RNA-editing factors is potentially longer than previously recognised. We used the redefined motifs to develop accurate and consistent annotations of PPR sequences from 109 genomes. We report a high error rate in PPR gene models in many public plant proteomes, due to gene fusions and insertions of spurious introns. These consistently annotated datasets across a wide range of species are valuable resources for future comparative genomics studies, and an essential pre-requisite for accurate large-scale computational predictions of PPR targets. We have created a web portal () that provides open access to these resources for the community.