A complete statistical model for calibration of RNA-seq counts using external spike-ins and maximum likelihood theory

A complete statistical model for calibration of RNA-seq counts using external spike-ins and maximum likelihood theory
复制标题

DOI:
10.1371/journal.pcbi.1006794
复制
发表时间:
2019-03-01
影响因子:
4.3
通讯作者:
Tranchina, Daniel
Tranchina, Daniel
中科院分区:
生物学2区
文献类型:
--
作者:
Athanasiadou, Rodoniki;Neymotin, Benjamin;Tranchina, Daniel

文献摘要

被引文献

相似文献

绝大多数高通量转录组分析的一个基本假设是,大多数基因的表达在样本中是不变的,并且细胞总RNA保持不变。然而,随着分析的实验系统数量的增加,不同的独立研究表明,这一假设经常被违反。我们提出了一种使用RNA尖刺的校准方法,可以测量RNA分子的绝对细胞丰度。我们将该方法应用于已知大小的细胞群中汇集的RNA。对于每个转录本,我们计算一个名义丰度,可以通过除以在单独实验中确定的比例因子转换为绝对丰度:转录本的产量系数相对于用相同方案测量的参考峰值。该方法是由最大似然理论,在一个完整的统计模型的背景下测序计数贡献的细胞RNA和刺入。计数是基于从固定数量的细胞中提取的样本,这些细胞中加入了固定数量的尖峰分子。我们通过应用两个全球表达数据集来说明和评估该方法,一个来自模型真核生物酿酒酵母,以不同的生长速度增殖,并在脊索动物中分化心咽部细胞系。我们在技术重复稀释研究和k倍验证研究中测试了该方法。作者总结:我们提出了一个完整的统计模型,用于分析来自细胞群体的RNA-seq数据,使用外部RNA刺入和最大似然方法来估计每个细胞的全基因组转录物。该模型包括细胞转录物数量的生物学变异性和采样噪声。我们通过简单地将计数乘以依赖于文库但与转录无关的比例因子,得出每个细胞转录本的无偏估计值。这是一个名义估计值,可以通过除以在单独实验中测量的转录本的相对产量系数转换为绝对估计值。负二项概率质量函数具有新颖的归一化(大小)因子,允许对有关每个转录本的绝对丰度对实验条件的依赖性的假设进行参数检验。我们的方法整合了所有重复和实验条件下每个RNA-seq实验的信息,以确定校准常数。我们用稀释研究和k倍交叉验证研究来检验该方法。我们通过应用酵母和海鞘的两个独立数据集来说明我们的方法,这些数据集是通过不同的文库制备协议获得的。我们表明,我们的方法检测全基因组的表达扩增,并将我们的方法与其他方法进行比较。
A fundamental assumption, common to the vast majority of high-throughput transcriptome analyses, is that the expression of most genes is unchanged among samples and that total cellular RNA remains constant. As the number of analyzed experimental systems increases however, different independent studies demonstrate that this assumption is often violated. We present a calibration method using RNA spike-ins that allows for the measurement of absolute cellular abundance of RNA molecules. We apply the method to pooled RNA from cell populations of known sizes. For each transcript, we compute a nominal abundance that can be converted to absolute by dividing by a scale factor determined in separate experiments: the yield coefficient of the transcript relative to that of a reference spike-in measured with the same protocol. The method is derived by maximum likelihood theory in the context of a complete statistical model for sequencing counts contributed by cellular RNA and spike-ins. The counts are based on a sample from a fixed number of cells to which a fixed population of spike-in molecules has been added. We illustrate and evaluate the method with applications to two global expression data sets, one from the model eukaryote Saccharomyces cerevisiae, proliferating at different growth rates, and differentiating cardiopharyngeal cell lineages in the chordate Ciona robusta. We tested the method in a technical replicate dilution study, and in a k-fold validation study.Author summary We present a complete statistical model for the analysis of RNA-seq data from a population of cells using external RNA spike-ins and a maximum-likelihood method for genome-wide estimation of transcripts per cell. The model includes biological variability of cellular transcript number and sampling noise. We derive an unbiased estimator of transcripts per cell for every transcript, given by simply multiplying the count by a library-dependent, but transcript-independent, scale factor. This is a nominal estimate that can be converted to an absolute estimate by dividing by the transcript's relative yield coefficient, measured in a separate experiment. A negative binomial probability mass function with novel normalization (size) factors allows for parametric testing of hypotheses concerning dependence of the absolute abundance of each transcript on experimental condition. Our method integrates information from every RNA-seq experiment across all replicates and experimental conditions to determine the calibration constants. We test the method with a dilution study and a k-fold cross-validation study. We illustrate our method with applications to two independent data sets from yeast and the sea squirt that were derived by different library preparation protocols. We show that our methods detect genome-wide amplification of expression, and we compare our method to others.