PRODUCTION OF A PRELIMINARY QUALITY CONTROL PIPELINE FOR SINGLE NUCLEI RNA-SEQ AND ITS APPLICATION IN THE ANALYSIS OF CELL TYPE DIVERSITY OF POST-MORTEM HUMAN BRAIN NEOCORTEX.

PRODUCTION OF A PRELIMINARY QUALITY CONTROL PIPELINE FOR SINGLE NUCLEI RNA-SEQ AND ITS APPLICATION IN THE ANALYSIS OF CELL TYPE DIVERSITY OF POST-MORTEM HUMAN BRAIN NEOCORTEX.
复制标题

DOI:
10.1142/9789813207813_0052
复制
发表时间:
2017
影响因子:
--
通讯作者:
Scheuermann RH
Scheuermann RH
中科院分区:
其他
文献类型:
--
作者:
Aevermann B;McCorrison J;Venepally P;Hodge R;Bakken T;Miller J;Novotny M;Tran DN;Diezfuertes F;Christiansen L;Zhang F;Steemers F;Lasken RS;Lein ED;Schork N;Scheuermann RH

文献摘要

被引文献

相似文献

下一代单细胞或单核RNA含量测序(sc/nRNA-seq)已成为了解多细胞生物和环境生态系统的细胞复杂性和多样性的有力方法。然而,该程序从相对少量的起始材料开始,从而推动了实验室程序所需的限制,这一事实表明,样品质量控制(QC)的谨慎方法对于减少下游分析应用中技术噪声和样品偏差的影响至关重要。在这里,我们提出了一个样本级质量控制的初步框架,该框架基于一系列定量实验室和数据指标的收集,这些指标被用作使用随机森林机器学习方法构建QC分类模型的特征。我们将这一初始框架应用于由2272个单核RNA-seq结果组成的数据集,并确定约79%的样本是高质量的。从下游分析中去除质量差的样品可以改善细胞类型聚类结果。此外,该方法还确定了与唯一或重复读取的比例以及经过质量修剪后剩余读取的比例相关的定量特征,作为合格/不合格分类的有用特征。构建和使用分类模型来识别低质量样本,为sc/ rna -seq质量控制提供了一种客观和可扩展的方法。
Next generation sequencing of the RNA content of single cells or single nuclei (sc/nRNA-seq) has become a powerful approach to understand the cellular complexity and diversity of multicellular organisms and environmental ecosystems. However, the fact that the procedure begins with a relatively small amount of starting material, thereby pushing the limits of the laboratory procedures required, dictates that careful approaches for sample quality control (QC) are essential to reduce the impact of technical noise and sample bias in downstream analysis applications. Here we present a preliminary framework for sample level quality control that is based on the collection of a series of quantitative laboratory and data metrics that are used as features for the construction of QC classification models using random forest machine learning approaches. We’ve applied this initial framework to a dataset comprised of 2272 single nuclei RNA-seq results and determined that ~79% of samples were of high quality. Removal of the poor quality samples from downstream analysis was found to improve the cell type clustering results. In addition, this approach identified quantitative features related to the proportion of unique or duplicate reads and the proportion of reads remaining after quality trimming as useful features for pass/fail classification. The construction and use of classification models for the identification of poor quality samples provides for an objective and scalable approach to sc/nRNA-seq quality control.