A revisit of RSEM generative model and its EM algorithm for quantifying transcript abundances

A revisit of RSEM generative model and its EM algorithm for quantifying transcript abundances
复制标题

重新审视用于量化转录本丰度的 RSEM 生成模型及其 EM 算法

DOI:
10.1101/503672
复制
发表时间:
2018
期刊:
bioRxiv
影响因子:
--
通讯作者:
Son K. Pham
Son K. Pham
中科院分区:
--
文献类型:
--
作者:
Hy Vuong;Thao T. Truong;Thang N Tran;Son K. Pham

文献摘要

被引文献

相似文献

RSEM主要因其在转录本丰度定量方面的准确性而闻名。然而,与最近的量化工具相比,它的量化时间非常长。本文对RSEM的EM算法进行了改进。特别是,我们得出准确的M步更新,以消除不正确的启发式更新RSEM。我们还实现了一些优化,将量化时间减少了约一百倍,同时与RSEM相比仍然具有更好的准确性。特别是,我们注意到不同的参数有不同的收敛速度,因此我们识别并删除早期收敛的参数,以显着降低模型的复杂性,在进一步的迭代,我们还使用SQUAREM方法,以进一步加快收敛速度。我们在一个名为Hera-EM的包中实现了这些修订,源代码可在https://github.com/bioturing/hera/tree/master/hera-EM获得
RSEM has been mainly known for its accuracy in transcript abundance quantification. However, its quantification time is extremely high compared to that of recent quantification tools. In this paper, we revised the RSEM’s EM algorithm. In particular, we derived accurate M-step updates to eliminate incorrect heuristic updates in RSEM. We also implement some optimizations that reduce the quantification time about a hundred times while still have better accuracy compared to RSEM. In particular, we noticed that different parameters have different convergence rates, therefore we identified and removed early converged parameters to significantly reduce the model complexity in further iterations, and we also use SQUAREM method to further speed up the convergence rate. We implemented these revisions in a packaged named Hera-EM, with source code available at: https://github.com/bioturing/hera/tree/master/hera-EM