Poly(A)-DG: A deep-learning-based domain generalization method to identify cross-species Poly(A) signal without prior knowledge from target species.
Poly(A)-DG: A deep-learning-based domain generalization method to identify cross-species Poly(A) signal without prior knowledge from target species.
复制标题
DOI:
10.1371/journal.pcbi.1008297
复制
发表时间:
2020-11
影响因子:
4.3
通讯作者:
Xu M
中科院分区:
文献类型:
--
作者:
Zheng Y;Wang H;Zhang Y;Gao X;Xing EP;Xu M
In eukaryotes, polyadenylation (poly(A)) is an essential process during mRNA maturation. Identifying the cis-determinants of poly(A) signal (PAS) on the DNA sequence is the key to understand the mechanism of translation regulation and mRNA metabolism. Although machine learning methods were widely used in computationally identifying PAS, the need for tremendous amounts of annotation data hinder applications of existing methods in species without experimental data on PAS. Therefore, cross-species PAS identification, which enables the possibility to predict PAS from untrained species, naturally becomes a promising direction. In our works, we propose a novel deep learning method named Poly(A)-DG for cross-species PAS identification. Poly(A)-DG consists of a Convolution Neural Network-Multilayer Perceptron (CNN-MLP) network and a domain generalization technique. It learns PAS patterns from the training species and identifies PAS in target species without re-training. To test our method, we use four species and build cross-species training sets with two of them and evaluate the performance of the remaining ones. Moreover, we test our method against insufficient data and imbalanced data issues and demonstrate that Poly(A)-DG not only outperforms state-of-the-art methods but also maintains relatively high accuracy when it comes to a smaller or imbalanced training set. The key to understanding the mechanism of translation regulation and mRNA metabolism is to identify the cis-determinants of PAS on the DNA sequence. PAS leads to correct identification of Poly(A) sites which play an essential role in understanding human diseases. While many researchers have employed deep learning methods to improve the performance of PAS identification, an underlying problem is the expensive and time-consuming nature of PAS data collection, which makes the application of deep learning models for identifying PAS from a broad range of species a tough task. We attempt to use domain generalization methods, inspired by its thrive in the field of computer vision, to overcome the insufficient annotation data challenge in PAS data. Here, empirical results suggest that our proposed model Poly(A)-DG can extract species-invariant features from multiple training species and be directly applied to the target species without fine-tuning. Furthermore, Poly(A)-DG is a promising practical tool for PAS identification with its stable performance on insufficient or species-imbalanced training data. We share the implementation of our proposed model on the GitHub. (https://github.com/Szym29/PolyADG).
登录
查看更多内容
DOI:
10.1093/bioinformatics/bts565
发表时间:
2012-12-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
通讯作者:
Li W
影响因子:
3.9
作者:
Gao, Xin;Zhang, Jie;Hakonarson, Hakon
通讯作者:
Hakonarson, Hakon
影响因子:
4.1
作者:
MacDonald, CC;Redondo, JL
通讯作者:
Redondo, JL
影响因子:
3
作者:
Ji, Guoli;Zheng, Jianti;Shen, Yingjia;Wu, Xiaohui;Jiang, Ronghan;Lin, Yun;Loke, Johnny C;Davis, Kimberly M;Reese, Greg J;Li, Qingshun Quinn
通讯作者:
Li, Qingshun Quinn
影响因子:
64.8
作者:
Jan, Calvin H.;Friedman, Robin C.;Ruby, J. Graham;Bartel, David P.
通讯作者:
Bartel, David P.