Deep learning for genomics using Janggu

Deep learning for genomics using Janggu
复制标题

DOI:
10.1038/s41467-020-17155-y
复制
发表时间:
2020-07-13
影响因子:
16.6
通讯作者:
Akalin, Altuna
Akalin, Altuna
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Kopp, Wolfgang;Monti, Remo;Akalin, Altuna

文献摘要

被引文献

相似文献

近年来,大量应用已经证明了深度学习在增进对生物过程的理解方面的潜力。然而,迄今为止开发的大多数深度学习工具都是为了解决固定数据集和/或固定模型架构上的特定问题而设计的。在这里,我们介绍 Janggu,一个 Python 库,有助于基因组学应用的深度学习,旨在简化数据采集和模型评估。其主要功能之一是特殊的数据集对象,它形成了一个统一且灵活的基因组数据采集和预处理框架,可以通过可重用的组件简化未来的研究应用程序。通过类似 numpy 的接口,这些数据集对象直接与流行的深度学习库兼容,包括 keras 或 pytorch。 Janggu 提供了将预测可视化为基因组轨迹或将其导出为 bigWig 格式以及基于 keras 模型的实用程序的可能性。我们在几个深度学习基因组学应用中说明了 Janggu 的功能。首先,我们评估了预测转录因子 JunD 结合位点任务的不同模型拓扑。其次,我们展示了用于预测染色质效应的已发表模型的框架。第三,我们表明可以使用 DNase 超敏性、组蛋白修饰和 DNA 序列特征来预测 CAGE 测量的启动子使用情况。由于 Janggu 的一项新颖功能允许我们包含高阶序列特征,因此我们提高了这些模型的性能。我们相信,Janggu 将有助于显着减少基因组学深度学习应用的重复编程开销,并使计算生物学家能够快速评估生物学假设。
In recent years, numerous applications have demonstrated the potential of deep learning for an improved understanding of biological processes. However, most deep learning tools developed so far are designed to address a specific question on a fixed dataset and/or by a fixed model architecture. Here we present Janggu, a python library facilitates deep learning for genomics applications, aiming to ease data acquisition and model evaluation. Among its key features are special dataset objects, which form a unified and flexible data acquisition and pre-processing framework for genomics data that enables streamlining of future research applications through reusable components. Through a numpy-like interface, these dataset objects are directly compatible with popular deep learning libraries, including keras or pytorch. Janggu offers the possibility to visualize predictions as genomic tracks or by exporting them to the bigWig format as well as utilities for keras-based models. We illustrate the functionality of Janggu on several deep learning genomics applications. First, we evaluate different model topologies for the task of predicting binding sites for the transcription factor JunD. Second, we demonstrate the framework on published models for predicting chromatin effects. Third, we show that promoter usage measured by CAGE can be predicted using DNase hypersensitivity, histone modifications and DNA sequence features. We improve the performance of these models due to a novel feature in Janggu that allows us to include high-order sequence features. We believe that Janggu will help to significantly reduce repetitive programming overhead for deep learning applications in genomics, and will enable computational biologists to rapidly assess biological hypotheses.