COSSMO: predicting competitive alternative splice site selection using deep learning

COSSMO: predicting competitive alternative splice site selection using deep learning
复制标题

DOI:
10.1093/bioinformatics/bty244
复制
发表时间:
2018-07-01
期刊:
影响因子:
5.8
通讯作者:
Frey, Brendan J.
Frey, Brendan J.
中科院分区:
生物学3区
文献类型:
--
作者:
Bretschneider, Hannes;Gandhi, Shreshth;Frey, Brendan J.

文献摘要

被引文献

相似文献

动机:选择性剪接位点的选择本质上是竞争性的,给定剪接位点被使用的可能性也取决于邻近位点的强度。在这里,我们提出了一个新的模型,称为竞争性剪接位点模型(COSSMO),它显式地解释了这些竞争效应,并预测了任何数量的假定剪接位点上的百分比选择指数(PSI)分布。我们将选择性剪接事件模拟为以固定的上游5‘供体位点为条件的3’受体位点的选择或以固定的3‘受体位点为条件的5’供体位点的选择。我们构建了四种不同的体系结构,分别使用卷积层、通信层、长短期记忆和残差网络来仅从序列中学习相关主题。我们还从基因组注释和RNA-Seq读取数据中构建了一个新的数据集,用于训练我们的模型。结果:COSSMO能够在未知测试数据上预测最频繁使用的剪接位点,准确率为70%,对PSI分布建模的R-2为0.6。我们可视化了COSSMO从序列中学习的基序,并表明COSSMO识别一致的剪接位点序列和许多已知的剪接因子,具有很高的特异性。
Motivation: Alternative splice site selection is inherently competitive and the probability of a given splice site to be used also depends on the strength of neighboring sites. Here, we present a new model named the competitive splice site model (COSSMO), which explicitly accounts for these competitive effects and predicts the percent selected index (PSI) distribution over any number of putative splice sites. We model an alternative splicing event as the choice of a 3' acceptor site conditional on a fixed upstream 5' donor site or the choice of a 5' donor site conditional on a fixed 3' acceptor site. We build four different architectures that use convolutional layers, communication layers, long short-term memory and residual networks, respectively, to learn relevant motifs from sequence alone. We also construct a new dataset from genome annotations and RNA-Seq read data that we use to train our model.Results: COSSMO is able to predict the most frequently used splice site with an accuracy of 70% on unseen test data, and achieve an R-2 of 0.6 in modeling the PSI distribution. We visualize the motifs that COSSMO learns from sequence and show that COSSMO recognizes the consensus splice site sequences and many known splicing factors with high specificity.