An Annotation Agnostic Algorithm for Detecting Nascent RNA Transcripts in GRO-Seq

An Annotation Agnostic Algorithm for Detecting Nascent RNA Transcripts in GRO-Seq
复制标题

DOI:
10.1109/tcbb.2016.2520919
复制
发表时间:
2017-09-01
影响因子:
4.5
通讯作者:
Dowell, Robin D.
Dowell, Robin D.
中科院分区:
工程技术3区
文献类型:
--
作者:
Azofeifa, Joseph G.;Allen, Mary A.;Dowell, Robin D.

文献摘要

被引文献

相似文献

我们提出了一种快速,简单的算法来检测新生RNA转录的全球核运行测序(GRO-seq)。GRO-seq是一种相对较新的方案,它从活跃的聚合酶中捕获新生转录物,提供对真实转录的直接读出。大多数传统的测定,如RNA-seq,测量受转录、转录后加工和RNA稳定性影响的稳态RNA水平。然而,GRO-seq数据提出了独特的分析挑战,这些挑战才刚刚开始得到解决。在这里,我们描述了一种新的算法,快速读取缝合器(FStitch),它利用了两种流行的机器学习技术,隐马尔可夫模型和逻辑回归,来分类基因组的哪些区域被转录。给定一个小的用户定义的训练集,我们的算法是准确的,对不同的读取深度,注释不可知,和快速的鲁棒性。在不需要事先注释的情况下对GRO-seq数据的分析揭示了对转录过程的几个方面的令人惊讶的新见解。
We present a fast and simple algorithm to detect nascent RNA transcription in global nuclear run-on sequencing (GRO-seq). GRO-seq is a relatively new protocol that captures nascent transcripts from actively engaged polymerase, providing a direct read-out on bona fide transcription. Most traditional assays, such as RNA-seq, measure steady state RNA levels which are affected by transcription, post-transcriptional processing, and RNA stability. GRO-seq data, however, presents unique analysis challenges that are only beginning to be addressed. Here, we describe a new algorithm, Fast Read Stitcher (FStitch), that takes advantage of two popular machine-learning techniques, hidden Markov models and logistic regression, to classify which regions of the genome are transcribed. Given a small user-defined training set, our algorithm is accurate, robust to varying read depth, annotation agnostic, and fast. Analysis of GRO-seq data without a priori need for annotation uncovers surprising new insights into several aspects of the transcription process.