DNABERT-S: LEARNING SPECIES-AWARE DNA EMBEDDING WITH GENOME FOUNDATION MODELS

DNABERT-S: LEARNING SPECIES-AWARE DNA EMBEDDING WITH GENOME FOUNDATION MODELS
复制标题

DOI:
10.48550/arxiv.2402.08777
复制
发表时间:
2024-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhihan Zhou;Weimin Wu;Harrison Ho;Jiayi Wang;Lizhen Shi;R. Davuluri;Zhong Wang;Han Liu
Zhihan Zhou;Weimin Wu;Harrison Ho;Jiayi Wang;Lizhen Shi;R. Davuluri;Zhong Wang;Han Liu
中科院分区:
其他
文献类型:
--
作者:
Zhihan Zhou;Weimin Wu;Harrison Ho;Jiayi Wang;Lizhen Shi;R. Davuluri;Zhong Wang;Han Liu

文献摘要

相似文献

有效的DNA嵌入在基因组分析中仍然至关重要,特别是在缺乏用于模型微调的标记数据的情况下,尽管基因组基础模型取得了重大进展。一个最好的例子是元基因组学,这是微生物组研究中的一个关键过程,目的是从可能来自数千个不同的、通常没有特征的物种的DNA序列的复杂混合物中,按物种对DNA序列进行分组。为了填补缺乏有效的DNA嵌入模型的不足,我们引入了DNABERT-S,这是一个专门创建物种感知DNA嵌入的基因组基础模型。为了鼓励对容易出错的长读DNA序列的有效嵌入,我们引入了流形实例混合(MI-MIX),这是一个对比目标,将DNA序列的隐藏表示混合在随机选择的层上,并训练模型在输出层识别和区分这些混合比例。我们通过拟议的课程对比学习(C2LR)策略进一步加强了这一点。在18个不同的数据集上的实证结果显示了DNABERT-S的显著表现。它在10枪物种分类中的表现优于顶级基线,只需2枪训练,同时在物种聚类中的调整兰德指数(ARI)翻了一番,并在元基因组学分类中显著增加了正确识别的物种数量。代码、数据和预先培训的模型可在https://github.com/Zhihan1996/DNABERT_S.上公开获得
Effective DNA embedding remains crucial in genomic analysis, particularly in scenarios lacking labeled data for model fine-tuning, despite the significant advancements in genome foundation models. A prime example is metagenomics binning, a critical process in microbiome research that aims to group DNA sequences by their species from a complex mixture of DNA sequences derived from potentially thousands of distinct, often uncharacterized species. To fill the lack of effective DNA embedding models, we introduce DNABERT-S, a genome foundation model that specializes in creating species-aware DNA embeddings. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 18 diverse datasets showed DNABERT-S’s remarkable performance. It outperforms the top baseline’s performance in 10-shot species classification with just a 2-shot training while doubling the Adjusted Rand Index (ARI) in species clustering and substantially increasing the number of correctly identified species in metagenomics binning. The code, data, and pre-trained model are publicly available at https://github.com/Zhihan1996/DNABERT_S.