FactorNet: A deep learning framework for predicting cell type specific transcription factor binding from nucleotide-resolution sequential data

FactorNet: A deep learning framework for predicting cell type specific transcription factor binding from nucleotide-resolution sequential data
复制标题

DOI:
10.1016/j.ymeth.2019.03.020
复制
发表时间:
2019-08-15
期刊:
影响因子:
4.8
通讯作者:
Xie, Xiaohui
Xie, Xiaohui
中科院分区:
生物学3区
文献类型:
--
作者:
Quang, Daniel;Xie, Xiaohui

文献摘要

被引文献

相似文献

由于存在大量转录因子 (TF) 和细胞类型,查询所有有效 TF/细胞类型对的结合谱在实验上是不可行的。为了解决这个问题,我们开发了一种称为 FactorNet 的卷积循环神经网络模型,用于通过计算来估算缺失的结合数据。 FactorNet 利用参考细胞类型的结合数据进行训练,利用各种特征(包括基因组序列、基因组注释、基因表达和信号数据(例如 DNase I 切割))对测试细胞类型进行预测。 FactorNet 实施了多种方便的策略来减少运行时间和内存消耗。通过可视化神经网络模型,我们可以解释模型如何预测结合。我们还研究了影响跨单元类型准确性的变量,并提供了改进该领域的建议。我们的方法在 ENCODE-DREAM 体内转录因子结合位点预测挑战赛中名列前茅,在 13 个最终轮评估 TF/细胞类型对中的 6 个中获得第一名,是所有参赛团队中最多的。 FactorNet 源代码是公开的,允许用户重现我们在 ENCODE-DREAM 挑战赛中的方法。
Due to the large numbers of transcription factors (TFs) and cell types, querying binding profiles of all valid TF/cell type pairs is not experimentally feasible. To address this issue, we developed a convolutional-recurrent neural network model, called FactorNet, to computationally impute the missing binding data. FactorNet trains on binding data from reference cell types to make predictions on testing cell types by leveraging a variety of features, including genomic sequences, genome annotations, gene expression, and signal data, such as DNase I cleavage. FactorNet implements several convenient strategies to reduce runtime and memory consumption. By visualizing the neural network models, we can interpret how the model predicts binding. We also investigate the variables that affect cross-cell type accuracy, and offer suggestions to improve upon this field. Our method ranked among the top teams in the ENCODE-DREAM in vivo Transcription Factor Binding Site Prediction Challenge, achieving first place on six of the 13 final round evaluation TF/cell type pairs, the most of any competing team. The FactorNet source code is publicly available, allowing users to reproduce our methodology from the ENCODE-DREAM Challenge.