A Distributed Semi-Supervised Platform for DNase-Seq Data Analytics using Deep Generative Convolutional Networks

A Distributed Semi-Supervised Platform for DNase-Seq Data Analytics using Deep Generative Convolutional Networks
复制标题

DOI:
10.1145/3233547.3233601
复制
发表时间:
2018-08
期刊:
Proceedings of the 2018 ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics
影响因子:
--
通讯作者:
Shayan Shams;Richard Platania;Joohyun Kim;Jian Zhang-;Kisung Lee;Seungwon Yang;Seung-Jong Park
Shayan Shams;Richard Platania;Joohyun Kim;Jian Zhang-;Kisung Lee;Seungwon Yang;Seung-Jong Park
中科院分区:
其他
文献类型:
--
作者:
Shayan Shams;Richard Platania;Joohyun Kim;Jian Zhang-;Kisung Lee;Seungwon Yang;Seung-Jong Park

文献摘要

被引文献

相似文献

提出了一种分析DNase-seq数据集的深度学习方法,该方法在揭示转录调控机制的生物学基础方面具有很好的潜力。对这些机制的进一步了解可以导致生命科学特别是药物、生物标记物的发现和癌症研究的重要进展。受最近深度学习领域显著进步的激励,我们开发了一个平台,深度半监督DNA酶序列分析(深度半监督DNA酶序列分析)。主要由深度生成卷积网络(ConvNets)支持,最显著的方面是半监督学习的能力,这对于经常被较少数量的标记数据困扰的常见生物学环境非常有益。此外,我们研究了一种基于k-mer的连续向量空间表示,试图通过考虑与相邻核苷酸之间基于位置的关系相关联的特征的生物序列的性质来进一步提高学习能力。该算法采用了一种改进的梯形网络作为生成模型的底层结构,并利用大规模DNase-seq实验中的序列在细胞类型分类任务中展示了其性能。在完全监督和半监督两种情况下,我们报告了该算法的性能。在完全监督和半监督两种情况下,DSSDA的分类准确率分别为94.6%和94.6%。在半监督情况下,即使在使用不到10%的标记数据的情况下,DSSDA的性能也与使用全数据集的其他ConvNet相当。我们的结果强调,为了应对具有挑战性的基因组序列数据集,需要一种更好的深度学习方法来学习潜在特征和表示。
A deep learning approach for analyzing DNase-seq datasets is presented, which has promising potentials for unraveling biological underpinnings on transcription regulation mechanisms. Further understanding of these mechanisms can lead to important advances in life sciences in general and drug, biomarker discovery, and cancer research in particular. Motivated by recent remarkable advances in the field of deep learning, we developed a platform, Deep Semi-Supervised DNase-seq Analytics (DSSDA). Primarily empowered by deep generative Convolutional Networks (ConvNets), the most notable aspect is the capability of semi-supervised learning, which is highly beneficial for common biological settings often plagued with a less sufficient number of labeled data. In addition, we investigated a k-mer based continuous vector space representation, attempting further improvement on learning power with the consideration of the nature of biological sequences for features associated with locality-based relationships between neighboring nucleotides. DSSDA employs a modified Ladder Network for underlying generative model architecture, and its performance is demonstrated on the cell type classification task using sequences from large-scale DNase-seq experiments. We report the performance of DSSDA in both fully-supervised setting, in which DSSDA outperforms widely-known ConvNet models (94.6% classification accuracy), and semi-supervised setting for which, even with less than 10% of labeled data, DSSDA performs relatively comparable to other ConvNets using the full data set. Our results underscore, in order to deal with challenging genomic sequence datasets, the need of a better deep learning method to learn latent features and representation.