Classifying next-generation sequencing data using a zero-inflated Poisson model

Classifying next-generation sequencing data using a zero-inflated Poisson model
复制标题

使用零膨胀泊松模型对下一代测序数据进行分类

DOI:
10.1093/bioinformatics/btx768
复制
发表时间:
2018-04-15
期刊:
影响因子:
5.8
通讯作者:
Tong, Tiejun
Tong, Tiejun
中科院分区:
生物学3区
文献类型:
--
作者:
Zhou, Yan;Wan, Xiang;Tong, Tiejun

文献摘要

被引文献

相似文献

动机:随着高通量技术的发展,RNA测序(RNA-seq)作为一种替代基因表达分析的方法越来越受欢迎,例如RNA图谱和分类。利用RNA-SEQ数据识别新患者属于哪种疾病已经被认为是医学研究中的一个重要问题。由于RNA-SEQ数据是离散的,为分类微阵列数据而开发的统计方法不能容易地应用于RNASEQ数据分类。Witten在2011年提出了泊松线性判别分析(PLDA)来对RNA-SEQ数据进行分类。然而,注意,计数数据集的特征经常是真实的RNA-seq或microRNA序列数据中的多余零(即,当序列深度不够或具有18-30个核苷酸长度的小RNA时)。因此,需要开发一种新的模型来分析多个零的RNA-seq数据。结果:本文提出了一种分析多个零的RNA-seq数据的零膨胀泊松Logistic判别分析(ZIPLDA)。新方法假定数据来自两种分布的混合:一种是零点质量,另一种服从泊松分布。然后,我们考虑了模型中观察到零的概率和基因的平均值以及测序深度之间的逻辑关系。仿真研究表明,所提出的方法在很大范围内都优于或至少与现有方法一样好。文中还分析了乳腺癌RNA-seq数据集和microRNA-seq数据集两个真实数据集,它们与我们提出的方法优于现有竞争对手的模拟结果相吻合。
Motivation: With the development of high-throughput techniques, RNA-sequencing (RNA-seq) is becoming increasingly popular as an alternative for gene expression analysis, such as RNAs profiling and classification. Identifying which type of diseases a new patient belongs to with RNA-seq data has been recognized as a vital problem in medical research. As RNA-seq data are discrete, statistical methods developed for classifying microarray data cannot be readily applied for RNAseq data classification. Witten proposed a Poisson linear discriminant analysis (PLDA) to classify the RNA-seq data in 2011. Note, however, that the count datasets are frequently characterized by excess zeros in real RNA-seq or microRNA sequence data (i.e. when the sequence depth is not enough or small RNAs with the length of 18-30 nucleotides). Therefore, it is desired to develop a new model to analyze RNA-seq data with an excess of zeros.Results: In this paper, we propose a Zero-Inflated Poisson Logistic Discriminant Analysis (ZIPLDA) for RNA-seq data with an excess of zeros. The new method assumes that the data are from a mixture of two distributions: one is a point mass at zero, and the other follows a Poisson distribution. We then consider a logistic relation between the probability of observing zeros and the mean of the genes and the sequencing depth in the model. Simulation studies show that the proposed method performs better than, or at least as well as, the existing methods in a wide range of settings. Two real datasets including a breast cancer RNA-seq dataset and a microRNA-seq dataset are also analyzed, and they coincide with the simulation results that our proposed method outperforms the existing competitors.