Poly(A)-DG: A deep-learning-based domain generalization method to identify cross-species Poly(A) signal without prior knowledge from target species.

Poly(A)-DG: A deep-learning-based domain generalization method to identify cross-species Poly(A) signal without prior knowledge from target species.
复制标题

DOI:
10.1371/journal.pcbi.1008297
复制
发表时间:
2020-11
影响因子:
4.3
通讯作者:
Xu M
Xu M
中科院分区:
生物学2区
文献类型:
--
作者:
Zheng Y;Wang H;Zhang Y;Gao X;Xing EP;Xu M

文献摘要

参考文献

被引文献

相似文献

在真核生物中,多聚腺苷酸化(poly(A))是mRNA成熟过程中必不可少的过程。鉴定DNA序列上多聚腺苷酸信号(poly(A)signal,PAS)的顺式决定簇是理解翻译调控和mRNA代谢机制的关键。虽然机器学习方法被广泛用于计算识别PAS,但对大量注释数据的需求阻碍了现有方法在没有PAS实验数据的物种中的应用。因此,跨物种PAS识别,这使得预测PAS从未经训练的物种的可能性,自然成为一个有前途的方向。在我们的工作中,我们提出了一种新的深度学习方法Poly(A)-DG,用于跨物种PAS识别。Poly(A)-DG由卷积神经网络-多层感知器(CNN-MLP)网络和域泛化技术组成。它从训练物种中学习PAS模式,并在不重新训练的情况下识别目标物种中的PAS。为了测试我们的方法,我们使用了四个物种,并使用其中两个构建了跨物种训练集,并评估了其余物种的性能。此外,我们针对数据不足和不平衡数据问题测试了我们的方法,并证明Poly(A)-DG不仅优于最先进的方法,而且在较小或不平衡的训练集上保持了相对较高的准确性。理解翻译调控和mRNA代谢机制的关键是确定PAS在DNA序列上的顺式决定簇。PAS导致正确识别Poly(A)位点,这些位点在理解人类疾病中发挥重要作用。虽然许多研究人员已经采用深度学习方法来提高PAS识别的性能,但一个潜在的问题是PAS数据收集的昂贵和耗时的性质,这使得应用深度学习模型从广泛的物种中识别PAS成为一项坚韧的任务。我们试图使用域泛化方法,其蓬勃发展的计算机视觉领域的启发,克服PAS数据中的注释数据不足的挑战。在这里,实证结果表明,我们提出的模型Poly(A)-DG可以从多个训练物种中提取物种不变的特征,并直接应用于目标物种而无需微调。此外,Poly(A)-DG在训练数据不足或物种不平衡的情况下具有稳定的性能,是一种很有前途的PAS识别实用工具。我们在GitHub上分享了我们提出的模型的实现。(https://github.com/Szym29/PolyADG).
In eukaryotes, polyadenylation (poly(A)) is an essential process during mRNA maturation. Identifying the cis-determinants of poly(A) signal (PAS) on the DNA sequence is the key to understand the mechanism of translation regulation and mRNA metabolism. Although machine learning methods were widely used in computationally identifying PAS, the need for tremendous amounts of annotation data hinder applications of existing methods in species without experimental data on PAS. Therefore, cross-species PAS identification, which enables the possibility to predict PAS from untrained species, naturally becomes a promising direction. In our works, we propose a novel deep learning method named Poly(A)-DG for cross-species PAS identification. Poly(A)-DG consists of a Convolution Neural Network-Multilayer Perceptron (CNN-MLP) network and a domain generalization technique. It learns PAS patterns from the training species and identifies PAS in target species without re-training. To test our method, we use four species and build cross-species training sets with two of them and evaluate the performance of the remaining ones. Moreover, we test our method against insufficient data and imbalanced data issues and demonstrate that Poly(A)-DG not only outperforms state-of-the-art methods but also maintains relatively high accuracy when it comes to a smaller or imbalanced training set. The key to understanding the mechanism of translation regulation and mRNA metabolism is to identify the cis-determinants of PAS on the DNA sequence. PAS leads to correct identification of Poly(A) sites which play an essential role in understanding human diseases. While many researchers have employed deep learning methods to improve the performance of PAS identification, an underlying problem is the expensive and time-consuming nature of PAS data collection, which makes the application of deep learning models for identifying PAS from a broad range of species a tough task. We attempt to use domain generalization methods, inspired by its thrive in the field of computer vision, to overcome the insufficient annotation data challenge in PAS data. Here, empirical results suggest that our proposed model Poly(A)-DG can extract species-invariant features from multiple training species and be directly applied to the target species without fine-tuning. Furthermore, Poly(A)-DG is a promising practical tool for PAS identification with its stable performance on insufficient or species-imbalanced training data. We share the implementation of our proposed model on the GitHub. (https://github.com/Szym29/PolyADG).
DOI: 10.1093/bioinformatics/bts565
发表时间: 2012-12-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
通讯作者: Li W
DOI: 10.1109/access.2018.2825996
发表时间: 2018-01-01
期刊: IEEE ACCESS
影响因子: 3.9
作者:
Gao, Xin;Zhang, Jie;Hakonarson, Hakon
通讯作者: Hakonarson, Hakon
DOI: 10.1016/s0303-7207(02)00044-8
发表时间: 2002-04-25
影响因子: 4.1
作者:
MacDonald, CC;Redondo, JL
通讯作者: Redondo, JL
植物信使RNA聚腺苷酸位点的预测建模。
DOI: 10.1186/1471-2105-8-43
发表时间: 2007-02-07
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Ji, Guoli;Zheng, Jianti;Shen, Yingjia;Wu, Xiaohui;Jiang, Ronghan;Lin, Yun;Loke, Johnny C;Davis, Kimberly M;Reese, Greg J;Li, Qingshun Quinn
通讯作者: Li, Qingshun Quinn
DOI: 10.1038/nature09616
发表时间: 2011-01-06
期刊: NATURE
影响因子: 64.8
作者:
Jan, Calvin H.;Friedman, Robin C.;Ruby, J. Graham;Bartel, David P.
通讯作者: Bartel, David P.