MSNet-4mC: learning effective multi-scale representations for identifying DNA N4-methylcytosine sites

MSNet-4mC: learning effective multi-scale representations for identifying DNA N4-methylcytosine sites
复制标题

DOI:
10.1093/bioinformatics/btac671
复制
发表时间:
2022-10-07
期刊:
影响因子:
5.8
通讯作者:
Akutsu, Tatsuya
Akutsu, Tatsuya
中科院分区:
生物学3区
文献类型:
--
作者:
Liu, Chunting;Song, Jiangning;Akutsu, Tatsuya

文献摘要

被引文献

相似文献

动机:N4-甲基胞嘧啶(4MC)是一种必不可少的表观遗传修饰,可调节广泛的生物学过程。但是,用于检测4MC位点的实验方法是耗时且劳动密集型的。作为替代方案,能够自动识别使用数据分析技术的4MC的计算方法成为合理的选择。一个主要的挑战是如何开发有效的方法来充分利用DNA序列中的复杂相互作用以提高预测能力。分子:在这项工作中,我们提出了MSNET-4MC,这是一种在卷积性运行的轻巧神经网络,具有多尺度的接受性的卷积操作在给定的DNA序列的短和长范围内感知跨元素关系的字段。考虑到不同物种中的候选人数量的强烈失衡,我们计算并应用了跨膜损失中的班级权重以平衡训练过程。广泛的基准测试实验表明,我们的方法可实现显着的性能提高,并且表现优于其他最先进的方法。
Motivation: N4-methylcytosine (4mC) is an essential kind of epigenetic modification that regulates a wide range of biological processes. However, experimental methods for detecting 4mC sites are time-consuming and labor-intensive. As an alternative, computational methods that are capable of automatically identifying 4mC with data analysis techniques become a reasonable option. A major challenge is how to develop effective methods to fully exploit the complex interactions within the DNA sequences to improve the predictive capability.Results: In this work, we propose MSNet-4mC, a lightweight neural network building upon convolutional operations with multi-scale receptive fields to perceive cross-element relationships over both short and long ranges of given DNA sequences. With strong imbalances in the number of candidates in different species in mind, we compute and apply class weights in the cross-entropy loss to balance the training process. Extensive benchmarking experiments show that our method achieves a significant performance improvement and outperforms other state-of-the-art methods.