featureCounts: an efficient general purpose program for assigning sequence reads to genomic features

featureCounts: an efficient general purpose program for assigning sequence reads to genomic features
复制标题

DOI:
10.1093/bioinformatics/btt656
复制
发表时间:
2014-04-01
期刊:
影响因子:
5.8
通讯作者:
Shi, Wei
Shi, Wei
中科院分区:
生物学3区
文献类型:
--
作者:
Liao, Yang;Smyth, Gordon K.;Shi, Wei

文献摘要

被引文献

相似文献

动机:下一代测序技术产生数百万个短序列,通常与参考基因组对齐。在许多应用中,下游分析所需的关键信息是每个基因组特征(例如每个外显子或每个基因)的读取数。对读进行计数的过程称为读汇总。阅读摘要是多种基因组分析所必需的,但迄今为止在文献中得到的关注相对较少。结果:我们提出了featurecots,一个适合从RNA或基因组DNA测序实验中产生的读取计数的读取摘要程序。featurecots实现了高效的染色体哈希和特征块技术。它比现有的方法要快得多(在基因级总结方面快了一个数量级),而且需要的计算机内存要少得多。它适用于单端或成对端读取,并为不同的测序应用提供了广泛的选择。
Motivation: Next-generation sequencing technologies generate millions of short sequence reads, which are usually aligned to a reference genome. In many applications, the key information required for downstream analysis is the number of reads mapping to each genomic feature, for example to each exon or each gene. The process of counting reads is called read summarization. Read summarization is required for a great variety of genomic analyses but has so far received relatively little attention in the literature.Results: We present featureCounts, a read summarization program suitable for counting reads generated from either RNA or genomic DNA sequencing experiments. featureCounts implements highly efficient chromosome hashing and feature blocking techniques. It is considerably faster than existing methods (by an order of magnitude for gene-level summarization) and requires far less computer memory. It works with either single or paired-end reads and provides a wide range of options appropriate for different sequencing applications.