Optimization of miRNA-seq data preprocessing.

Optimization of miRNA-seq data preprocessing.
复制标题

DOI:
10.1093/bib/bbv019
复制
发表时间:
2015-11
影响因子:
9.5
通讯作者:
McPherson JD
McPherson JD
中科院分区:
生物学2区
文献类型:
--
作者:
Tam S;Tsao MS;McPherson JD

文献摘要

被引文献

相似文献

过去二十年的microRNA(miRNA)研究已经巩固了这些小的非编码RNA作为许多生物过程的关键调节因子和有前途的疾病生物标志物的作用。高通量分析技术的同步发展进一步推进了我们对它们在全球范围内失调的影响的理解。目前,下一代测序是发现和定量miRNA的首选平台。尽管如此,对于在进行下游分析之前应如何预处理数据,目前还没有明确的共识。数据预处理是数据分析中的一个重要步骤:不可靠特征和噪声的存在会影响下游分析得出的结论。使用加标稀释研究,我们评估了几种通用比对器(BWA,Bowtie,Bowtie 2和Novoalign)和归一化方法(每百万计数,总计数缩放,上四分位数缩放,M的修剪平均值,DESeq,线性回归,循环黄土和分位数)对最终miRNA计数数据分布,方差,偏倚和差异表达分析准确性的影响。我们对小RNA测序实验中miRNA计数数据的提取和解释的最佳预处理方法提出了实用的建议。
The past two decades of microRNA (miRNA) research has solidified the role of these small non-coding RNAs as key regulators of many biological processes and promising biomarkers for disease. The concurrent development in high-throughput profiling technology has further advanced our understanding of the impact of their dysregulation on a global scale. Currently, next-generation sequencing is the platform of choice for the discovery and quantification of miRNAs. Despite this, there is no clear consensus on how the data should be preprocessed before conducting downstream analyses. Often overlooked, data preprocessing is an essential step in data analysis: the presence of unreliable features and noise can affect the conclusions drawn from downstream analyses. Using a spike-in dilution study, we evaluated the effects of several general-purpose aligners (BWA, Bowtie, Bowtie 2 and Novoalign), and normalization methods (counts-per-million, total count scaling, upper quartile scaling, Trimmed Mean of M, DESeq, linear regression, cyclic loess and quantile) with respect to the final miRNA count data distribution, variance, bias and accuracy of differential expression analysis. We make practical recommendations on the optimal preprocessing methods for the extraction and interpretation of miRNA count data from small RNA-sequencing experiments.