Empirical null estimation using zero-inflated discrete mixture distributions and its application to protein domain data.

Empirical null estimation using zero-inflated discrete mixture distributions and its application to protein domain data.
复制标题

DOI:
10.1111/biom.12779
复制
发表时间:
2018-06
期刊:
影响因子:
1.9
通讯作者:
Spouge JL
Spouge JL
中科院分区:
数学3区
文献类型:
--
作者:
Gauran IIM;Park J;Lim J;Park D;Zylstra J;Peterson T;Kann M;Spouge JL

文献摘要

参考文献

被引文献

相似文献

在最近的突变研究中,基于蛋白质结构域位置的分析比以基因为中心的方法越来越受欢迎,因为后者在考虑突变位置提供的功能背景方面存在局限性。这提出了一个大规模的同时推理问题,同时需要考虑数百个假设检验。本文旨在通过错误发现率(FDR)程序选择显著突变计数,同时控制给定水平的I型错误。一个主要假设是,突变计数遵循零膨胀模型,以便考虑计数模型中的真零和多余零。所考虑的一类模型是零膨胀广义泊松(ZIGP)分布。此外,我们假设存在一个截止值,使得从零分布生成的计数小于此值。我们提出了几种依赖于数据的方法来确定截止值。我们还考虑了基于筛选过程的两阶段程序,以便超过一定值的突变数量应被视为显著突变。模拟和蛋白质结构域数据集用于说明使用离散分布的混合估计经验零值的这一过程。总体而言,在保持对FDR的控制的同时,所提出的两阶段测试程序具有优越的经验力量。
In recent mutation studies, analyses based on protein domain positions are gaining popularity over gene-centric approaches since the latter have limitations in considering the functional context that the position of the mutation provides. This presents a large-scale simultaneous inference problem, with hundreds of hypothesis tests to consider at the same time. This paper aims to select significant mutation counts while controlling a given level of Type I error via False Discovery Rate (FDR) procedures. One main assumption is that the mutation counts follow a zero-inflated model in order to account for the true zeros in the count model and the excess zeros. The class of models considered is the Zero-inflated Generalized Poisson (ZIGP) distribution. Furthermore, we assumed that there exists a cut-off value such that smaller counts than this value are generated from the null distribution. We present several data-dependent methods to determine the cut-off value. We also consider a two-stage procedure based on screening process so that the number of mutations exceeding a certain value should be considered as significant mutations. Simulated and protein domain data sets are used to illustrate this procedure in estimation of the empirical null using a mixture of discrete distributions. Overall, while maintaining control of the FDR, the proposed two-stage testing procedure has superior empirical power.
DOI: 10.1038/onc.2008.343
发表时间: 2008-11-24
期刊: ONCOGENE
影响因子: 8
作者:
Jeanes, A.;Gottardi, C. J.;Yap, A. S.
通讯作者: Yap, A. S.
DOI: 10.1111/1467-9868.00346
发表时间: 2002-01-01
影响因子: 5.8
作者:
Storey, JD
通讯作者: Storey, JD
DOI: 10.1093/bioinformatics/btq447
发表时间: 2010-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Peterson, Thomas A.;Adadey, Asa;Kann, Maricel G.
通讯作者: Kann, Maricel G.
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1198/016214501753382129
发表时间: 2001-12-01
影响因子: 3.7
作者:
Efron, B;Tibshirani, R;Tusher, V
通讯作者: Tusher, V