Universal count correction for high-throughput sequencing.
Universal count correction for high-throughput sequencing.
复制标题
DOI:
10.1371/journal.pcbi.1003494
复制
发表时间:
2014-03
影响因子:
4.3
通讯作者:
Gifford DK
中科院分区:
文献类型:
--
作者:
Hashimoto TB;Edwards MD;Gifford DK
We show that existing RNA-seq, DNase-seq, and ChIP-seq data exhibit overdispersed per-base read count distributions that are not matched to existing computational method assumptions. To compensate for this overdispersion we introduce a nonparametric and universal method for processing per-base sequencing read count data called Fixseq. We demonstrate that Fixseq substantially improves the performance of existing RNA-seq, DNase-seq, and ChIP-seq analysis tools when compared with existing alternatives. High-throughput DNA sequencing has been adapted to measure diverse biological state information including RNA expression, chromatin accessibility, and transcription factor binding to the genome. The accurate inference of biological mechanism from sequence counts requires a model of how sequence counts are distributed. We show that presently used sequence count distribution models are typically inaccurate and present a new method called Fixseq to process counts to more closely follow existing count models. On typical datasets Fixseq improves the performance of existing tools for RNA-seq, DNase-seq, and ChIP-seq, while yielding complementary additional gains in cases where domain-specific tools are available.
登录
查看更多内容
影响因子:
46.9
作者:
Ji, Hongkai;Jiang, Hui;Ma, Wenxiu;Johnson, David S.;Myers, Richard M.;Wong, Wing H.
通讯作者:
Wong, Wing H.
影响因子:
14.9
作者:
Hansen KD;Brenner SE;Dudoit S
通讯作者:
Dudoit S
影响因子:
5.8
作者:
Li, Wei;Jiang, Tao
通讯作者:
Jiang, Tao
DOI:
10.1093/bioinformatics/bts055
发表时间:
2012-04-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
Jones DC;Ruzzo WL;Peng X;Katze MG
通讯作者:
Katze MG
影响因子:
4.3
作者:
Guo Y;Mahony S;Gifford DK
通讯作者:
Gifford DK