Benchmark analysis of algorithms for determining and quantifying full-length mRNA splice forms from RNA-seq data.

Benchmark analysis of algorithms for determining and quantifying full-length mRNA splice forms from RNA-seq data.
复制标题

DOI:
10.1093/bioinformatics/btv488
复制
发表时间:
2015-12-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Grant GR
Grant GR
中科院分区:
其他
文献类型:
--
作者:
Hayer KE;Pizarro A;Lahens NF;Hogenesch JB;Grant GR

文献摘要

被引文献

相似文献

动机:由于RNA测序(RNA- seq)相对于微阵列的优势,它在高度平行的基因表达分析中越来越受欢迎。例如,RNA-Seq有望提供全长剪接形式的准确鉴定和定量。为此目的已经开发了一些信息学包,但从原则上讲,短的读取使它成为一个难题。测序错误和多态性增加了进一步的复杂性。有必要进行研究以确定哪些算法表现最好,哪些算法表现足够。然而,缺乏独立和公正的基准研究。在这里,我们采用一种方法,同时使用模拟和实验基准数据来评估其准确性。结果:我们得出结论,即使使用理想的数据,大多数方法也是不准确的,一旦存在多种剪接形式、多态性、内含子信号、测序错误、比对错误、注释错误和其他复杂因素,没有一种方法是高度准确的。这些结果表明迫切需要进一步的算法开发。可用性和实施:模拟数据集和其他支持信息可在http://bioinf.itmat.upenn.edu/BEERS/bp2上找到补充信息:补充数据可在Bioinformatics在线上获得。联系:hayer@upenn.edu
Motivation: Because of the advantages of RNA sequencing (RNA-Seq) over microarrays, it is gaining widespread popularity for highly parallel gene expression analysis. For example, RNA-Seq is expected to be able to provide accurate identification and quantification of full-length splice forms. A number of informatics packages have been developed for this purpose, but short reads make it a difficult problem in principle. Sequencing error and polymorphisms add further complications. It has become necessary to perform studies to determine which algorithms perform best and which if any algorithms perform adequately. However, there is a dearth of independent and unbiased benchmarking studies. Here we take an approach using both simulated and experimental benchmark data to evaluate their accuracy. Results: We conclude that most methods are inaccurate even using idealized data, and that no method is highly accurate once multiple splice forms, polymorphisms, intron signal, sequencing errors, alignment errors, annotation errors and other complicating factors are present. These results point to the pressing need for further algorithm development. Availability and implementation: Simulated datasets and other supporting information can be found at http://bioinf.itmat.upenn.edu/BEERS/bp2 Supplementary information: Supplementary data are available at Bioinformatics online. Contact: hayer@upenn.edu