EmptyNN: A neural network based on positive and unlabeled learning to remove cell-free droplets and recover lost cells in scRNA-seq data.

EmptyNN: A neural network based on positive and unlabeled learning to remove cell-free droplets and recover lost cells in scRNA-seq data.
复制标题

DOI:
10.1016/j.patter.2021.100311
复制
发表时间:
2021-08-13
期刊:
Patterns (New York, N.Y.)
影响因子:
--
通讯作者:
Simon LM
Simon LM
中科院分区:
其他
文献类型:
--
作者:
Yan F;Zhao Z;Simon LM

文献摘要

参考文献

被引文献

相似文献

基于液滴的单细胞RNA测序(scRNA-seq)显著增加了每个实验的细胞数量,并彻底改变了个体转录组的研究。然而,为了最大限度地提高生物信号,需要强大的计算方法来区分无细胞和含细胞的液滴。在这里,我们引入了一种新的称为EmptyNN的细胞调用算法,该算法基于正无标签学习训练神经网络,以改进条形码的过滤。出于基准测试的目的,我们利用细胞散列和遗传变异来提供基本事实。EmptyNN在恢复丢失的细胞簇的同时准确地去除了游离的细胞液滴,在接收器工作特性下的面积分别达到了94.73%和96.30%。与当前最先进的蜂窝呼叫算法进行比较,证明了EmptyNN的优越性能。EmptyNN进一步应用于单核RNA测序(snRNA-seq)数据集,显示出良好的性能。因此,EmptyNN是增强scRNA-seq和snRNA-seq质量控制分析的有力工具。新的细胞调用算法EmptyNN提高了scRNA-seq数据集的质量EmptyNN准确地去除无细胞液滴并恢复真实细胞基准分析利用细胞哈希信息和遗传变异在细胞水平上以高通量测量基因表达的进展是由基于液滴的单细胞RNA测序(scRNA-seq)平台的出现推动的。基于液滴的scRNA-seq平台在每次实验中分析了大量细胞,加速了我们对生物学的理解。准确分类无细胞和含细胞液滴将最大限度地提高生物信号,便于下游分析。在这里,我们提出了一种新的称为EmptyNN的细胞调用算法,该算法基于正无标签学习训练神经网络,以改进条形码的过滤。我们的研究结果表明,EmptyNN优于现有的细胞调用方法,因此代表了一个强大的工具,以增强scRNA-seq和单核RNA测序质量控制分析。为了测量单个细胞的基因表达水平,使用基于液滴的单细胞RNA测序平台将细胞分离成油滴。然而,在硅分离得到的表达数据中的空液滴和含细胞液滴是具有挑战性的。我们的算法,称为EmptyNN,通过利用神经网络和未标记的积极学习,提高了空液滴和含细胞液滴的区别。
Droplet-based single-cell RNA sequencing (scRNA-seq) has significantly increased the number of cells profiled per experiment and revolutionized the study of individual transcriptomes. However, to maximize the biological signal, robust computational methods are needed to distinguish cell-free from cell-containing droplets. Here, we introduce a novel cell-calling algorithm called EmptyNN, which trains a neural network based on positive-unlabeled learning for improved filtering of barcodes. For benchmarking purposes, we leveraged cell hashing and genetic variation to provide ground truth. EmptyNN accurately removed cell-free droplets while recovering lost cell clusters, and achieved an area under the receiver operating characteristics of 94.73% and 96.30%, respectively. Comparisons to current state-of-the-art cell-calling algorithms demonstrated the superior performance of EmptyNN. EmptyNN was further applied to a single-nucleus RNA sequencing (snRNA-seq) dataset and showed good performance. Therefore, EmptyNN represents a powerful tool to enhance both scRNA-seq and snRNA-seq quality control analyses. The novel cell-calling algorithm EmptyNN improves the quality of scRNA-seq datasets EmptyNN accurately removes cell-free droplets and recovers genuine cells Benchmarking analyses leverage cell hashing information and genetic variation Advances in measuring gene expression at the cellular level at high throughput have been fueled by the advent of droplet-based single-cell RNA sequencing (scRNA-seq) platforms. Droplet-based scRNA-seq platforms profile a large number of cells per experiment and accelerate our understanding of biology. Accurate classification of cell-free and cell-containing droplets will maximize biological signal and facilitate downstream analysis. Here, we present a novel cell-calling algorithm called EmptyNN, which trains a neural network based on positive-unlabeled learning for improved filtering of barcodes. Our results indicate that EmptyNN outperforms existing cell-calling methods and, thus, represents a powerful tool to enhance both scRNA-seq and single-nucleus RNA sequencing quality control analyses. To measure the gene expression levels of an individual cell, cells are isolated into oil droplets using droplet-based single-cell RNA sequencing platforms. However, in silico separation of empty and cell-containing droplets in the resulting expression data is challenging. Our algorithm, called EmptyNN, improves distinction between empty and cell-containing droplets by leveraging neural networks and unlabeled-positive learning.
DOI: 10.1093/gigascience/giaa122
发表时间: 2020-12-10
期刊: GigaScience
影响因子: 9.2
作者:
Simon LM;Yan F;Zhao Z
通讯作者: Zhao Z
DOI: 10.1038/ncomms14049
发表时间: 2017-01-16
影响因子: 16.6
作者:
Zheng GX;Terry JM;Belgrader P;Ryvkin P;Bent ZW;Wilson R;Ziraldo SB;Wheeler TD;McDermott GP;Zhu J;Gregory MT;Shuga J;Montesclaros L;Underwood JG;Masquelier DA;Nishimura SY;Schnall-Levin M;Wyatt PW;Hindson CM;Bharadwaj R;Wong A;Ness KD;Beppu LW;Deeg HJ;McFarland C;Loeb KR;Valente WJ;Ericson NG;Stevens EA;Radich JP;Mikkelsen TS;Hindson BJ;Bielas JH
通讯作者: Bielas JH
DOI: 10.1038/s41598-020-67513-5
发表时间: 2020-07-03
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
Alvarez, Marcus;Rahmani, Elior;Pajukanta, Paivi
通讯作者: Pajukanta, Paivi
DOI: 10.1007/11564096_24
发表时间: 2005-01-01
期刊: MACHINE LEARNING: ECML 2005, PROCEEDINGS
影响因子: --
作者:
Li, XL;Liu, B
通讯作者: Liu, B
DOI: 10.1038/nmeth.4407
发表时间: 2017-10
期刊: Nature methods
影响因子: 48
作者:
Habib N;Avraham-Davidi I;Basu A;Burks T;Shekhar K;Hofree M;Choudhury SR;Aguet F;Gelfand E;Ardlie K;Weitz DA;Rozenblatt-Rosen O;Zhang F;Regev A
通讯作者: Regev A