SCOPE++: sequence classification of homoPolymer emissions.

SCOPE++: sequence classification of homoPolymer emissions.
复制标题

DOI:
10.1016/j.ygeno.2014.07.005
复制
发表时间:
2014-09
期刊:
影响因子:
4.4
通讯作者:
Karro, John E.
Karro, John E.
中科院分区:
生物学3区
文献类型:
--
作者:
Morton, James T.;Abrudan, Patricia;Figueroa, Nathanial;Liang, Chun;Karro, John E.

文献摘要

参考文献

相似文献

mRNA 多腺苷酸化(即在前体 mRNA 的 3' 端添加聚腺苷酸尾)是真核生物基因表达和调控的关键过程。为了了解控制聚腺苷酸化和其他相关生物过程的分子机制,在转录组测序数据中准确识别这些聚腺苷酸尾并将其与测序过程中添加的人工接头序列区分开来非常重要。但由于存在测序错误和转录后修饰,这些尾部的注释变得复杂。虽然确定给定转录片段中是否存在尾部是很简单的,但这些混淆使边界识别问题成为一个挑战;传统的种子和延伸算法很难准确识别这些 Poly(A) 尾部端点。此外,我们所知的所有现有工具都只专注于聚腺苷酸尾的修剪,未能提供研究聚腺苷酸化过程所需的详细信息。我们创建了 SCOPE++,这是一种在原始 mRNA 序列读数中查找 Poly(A) 尾部和其他同聚物的精确边界的工具。基于隐马尔可夫模型 (HMM) 方法,SCOPE++ 以适合大型序列集的速度准确识别容易出错的 EST/cDNA 数据或 RNA-Seq 数据中的特定均聚物序列。我们证明,我们的工具可以以高通量应用所需的速度以近乎完美的精度精确识别聚腺苷酸尾,为聚腺苷酸化研究提供宝贵的资源。
mRNA polyadenylation, the addition of a poly(A) tail to the 3'-end of pre-mRNA, is a process critical to gene expression and regulation in eukaryotes. To understand the molecular mechanisms governing polyadenylation and other relevant biological processes, it is important to identify these poly(A) tails accurately in transcriptome sequencing data and differentiate them from artificial adapter sequences added in the sequencing process. But the annotation of these tails is complicated by the presence of sequencing errors and post-transcriptional modifications. While determining that a tail is present in a given transcript fragment is straight-forward, these obfuscations make the problem of boundary identification a challenge; conventional seed-and-extend algorithms struggle to accurately identify these poly(A) tail end-points. Further, all existing tools that we are aware of focus exclusively on the trimming of poly(A) tails, failing to provide the detailed information needed for studying the polyadenylation process. We have created SCOPE++, a tool for finding the precise border of poly(A) tails and other homopolymers in raw mRNA sequence reads. Based on a Hidden Markov Model (HMM) approach, SCOPE++ accurately identifies specific homopolymer sequences in error-prone EST/cDNA data or RNA-Seq data at a speed appropriate for large sequence sets. We demonstrate that our tool can precisely identify poly(A) tails with near perfect accuracy at the speed required for high-throughput applications, providing a valuable resource for polyadenylation research.
DOI: 10.1186/1471-2105-11-38
发表时间: 2010-01-20
期刊: BMC bioinformatics
影响因子: 3
作者:
Falgueras J;Lara AJ;Fernández-Pozo N;Cantón FR;Pérez-Trabado G;Claros MG
通讯作者: Claros MG
DOI: 10.1109/massp.1986.1165342
发表时间: 2007-06-01
影响因子: --
作者:
Schuster-Bockler, Benjamin;Bateman, Alex
通讯作者: Bateman, Alex
DOI: 10.1016/j.cell.2009.06.016
发表时间: 2009-08-21
期刊: Cell
影响因子: 64.5
作者:
Mayr C;Bartel DP
通讯作者: Bartel DP
DOI: 10.1016/j.cell.2010.11.020
发表时间: 2010-12-10
期刊: Cell
影响因子: 64.5
作者:
Ozsolak F;Kapranov P;Foissac S;Kim SW;Fishilevich E;Monaghan AP;John B;Milos PM
通讯作者: Milos PM
DOI: 10.1261/rna.7610404
发表时间: 2004-11-01
期刊: RNA
影响因子: 4.5
作者:
Jin, YF;Bian, TF
通讯作者: Bian, TF