Bias Correction in RNA-Seq Short-Read Counts Using Penalized Regression

Bias Correction in RNA-Seq Short-Read Counts Using Penalized Regression
复制标题

使用惩罚回归校正 RNA-Seq 短读计数中的偏差

DOI:
--
复制
发表时间:
2012
影响因子:
1
通讯作者:
P. Ma
P. Ma
中科院分区:
--
文献类型:
--
作者:
David Dalpiaz;Xuming He;P. Ma

文献摘要

被引文献

相似文献

RNA-Seq产生数千万个短读段。当映射到基因组和/或参考转录物时,RNA-Seq数据可以通过非常大量的短读段计数来总结。准确的转录本定量,如基因表达计算,依赖于RNA-Seq短读计数中序列偏差的正确校正。我们使用一个线性模型的序列偏差,这是更灵活的比流行的泊松模型。我们使用惩罚回归方法拟合模型,该方法允许显着的降维。该算法可扩展用于建模RNA-Seq数据。我们证明了我们所提出的方法的优良性能,将其应用到真实的例子。这些方法是在开源代码中实现的,可以在R包lmbc中找到。
RNA-Seq produces tens of millions of short reads. When mapped to the genome and/or to the reference transcripts, RNA-Seq data can be summarized by a very large number of short-read counts. Accurate transcript quantification, such as gene expression calculation, relies on proper correction of sequence bias in the RNA-Seq short-read counts. We use a linear model for the sequence bias, which is much more flexible than the popular Poisson model. We fit the model using a penalized regression method, which allows for a significant dimension reduction. The algorithm is scalable for modeling RNA-Seq data. We demonstrate the excellent performance of our proposed method by applying it to real examples. The methods are implemented in open-source code, which is available in the R package lmbc.