Predicting RNA-protein binding sites and motifs through combining local and global deep convolutional neural networks

Predicting RNA-protein binding sites and motifs through combining local and global deep convolutional neural networks
复制标题

通过结合局部和全局深度卷积神经网络来预测 RNA-蛋白质结合位点和基序。

DOI:
10.1093/bioinformatics/bty364
复制
发表时间:
2018-10-15
期刊:
影响因子:
5.8
通讯作者:
Shen, Hong-Bin
Shen, Hong-Bin
中科院分区:
生物学3区
文献类型:
--
作者:
Pan, Xiaoyong;Shen, Hong-Bin

文献摘要

被引文献

相似文献

动机:RNA结合蛋白(RNA binding proteins,RBP)占真核生物蛋白质组的5-10%,在基因调控等生物学过程中发挥重要作用。RBP结合位点的实验检测仍然是耗时且高成本的。相反,使用从现有注释知识中学习的模式来计算预测RBP结合位点是一种快速方法。从生物学的角度来看,由局部序列衍生的局部结构上下文将被特定的RBP识别。然而,在使用深度学习的计算建模中,据我们所知,仅采用整个RNA序列的全局表示。到目前为止,在深度模型构建过程中忽略了局部序列信息。结果:在这项研究中,我们提出了一种计算方法iDeepE,通过结合全局和局部卷积神经网络(CNN),从RNA序列中预测RNA-蛋白质结合位点。对于全局CNN,我们将RNA序列填充到相同的长度。对于局部CNN,我们将RNA序列分割成多个重叠的固定长度序列,其中每个子序列是整个序列的信号通道。接下来,我们分别为多个序列和填充序列训练深度CNN来学习高级特征。最后,将本地和全局CNN的输出组合起来以改进预测。iDeepE在两个来自CLIP-seq的大规模数据集上表现出比最先进的方法更好的性能。我们还发现,当使用GPU时,本地CNN的运行速度比全局CNN快1.8倍,性能相当。我们的研究结果表明,iDeepE已经捕获了实验验证的结合mitifs。
Motivation: RNA-binding proteins (RBPs) take over 5-10% of the eukaryotic proteome and play key roles in many biological processes, e.g. gene regulation. Experimental detection of RBP binding sites is still time-intensive and high-costly. Instead, computational prediction of the RBP binding sites using patterns learned from existing annotation knowledge is a fast approach. From the biological point of view, the local structure context derived from local sequences will be recognized by specific RBPs. However, in computational modeling using deep learning, to our best knowledge, only global representations of entire RNA sequences are employed. So far, the local sequence information is ignored in the deep model construction process.Results: In this study, we present a computational method iDeepE to predict RNA-protein binding sites from RNA sequences by combining global and local convolutional neural networks (CNNs). For the global CNN, we pad the RNA sequences into the same length. For the local CNN, we split a RNA sequence into multiple overlapping fixed-length subsequences, where each subsequence is a signal channel of the whole sequence. Next, we train deep CNNs for multiple subsequences and the padded sequences to learn high-level features, respectively. Finally, the outputs from local and global CNNs are combined to improve the prediction. iDeepE demonstrates a better performance over state-of-the-art methods on two large-scale datasets derived from CLIP-seq. We also find that the local CNN runs 1.8 times faster than the global CNN with comparable performance when using GPUs. Our results show that iDeepE has captured experimentally verified binding mitifs.