DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome

DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome
复制标题

DOI:
10.1093/bioinformatics/btab083
复制
发表时间:
2021-02-04
期刊:
影响因子:
5.8
通讯作者:
Davuluri, Ramana, V
Davuluri, Ramana, V
中科院分区:
生物学3区
文献类型:
--
作者:
Ji, Yanrong;Zhou, Zhihan;Davuluri, Ramana, V

文献摘要

被引文献

相似文献

动机:破译非编码DNA的语言是基因组研究的基本问题之一。结果:针对这一挑战,提出了一种新的预先训练好的双向编码法DNABERT,以获取基因组DNA序列的全局性和可转移性。我们将DNABERT与最广泛使用的全基因组调控元件预测程序进行了比较,并证明了它的易用性、准确性和效率。我们表明,在使用小的特定于任务的标记数据进行容易的微调之后,单个预训练的转换器模型可以同时在预测启动子、剪接位点和转录因子结合位点方面获得最先进的性能。此外,DNABERT能够直接可视化输入序列中核苷酸水平的重要性和语义关系,以便更好地解释和准确识别保守的序列基序和功能遗传变异候选。最后,我们展示了预先训练的具有人类基因组的DNABERT甚至可以很容易地应用于其他具有优异性能的生物。我们预计,预先训练的DNABERT模型可以微调到许多其他序列分析任务。
Motivation: Deciphering the language of non-coding DNA is one of the fundamental problems in genome research. Gene regulatory code is highly complex due to the existence of polysemy and distant semantic relationship, which previous informatics methods often fail to capture especially in data-scarce scenarios.Results: To address this challenge, we developed a novel pre-trained bidirectional encoder representation, named DNABERT, to capture global and transferrable understanding of genomic DNA sequences based on up and downstream nucleotide contexts. We compared DNABERT to the most widely used programs for genome-wide regulatory elements prediction and demonstrate its ease of use, accuracy and efficiency. We show that the single pre-trained transformers model can simultaneously achieve state-of-the-art performance on prediction of promoters, splice sites and transcription factor binding sites, after easy fine-tuning using small task-specific labeled data. Further, DNABERT enables direct visualization of nucleotide-level importance and semantic relationship within input sequences for better interpretability and accurate identification of conserved sequence motifs and functional genetic variant candidates. Finally, we demonstrate that pre-trained DNABERT with human genome can even be readily applied to other organisms with exceptional performance. We anticipate that the pre-trained DNABERT model can be fined tuned to many other sequence analyses tasks.