Flexible bivariate correlated count data regression

Flexible bivariate correlated count data regression
复制标题

DOI:
10.1002/sim.8676
复制
发表时间:
2020-08-04
影响因子:
2
通讯作者:
Ho, Yen-Yi
Ho, Yen-Yi
中科院分区:
医学3区
文献类型:
--
作者:
Ma, Zichen;Hanson, Timothy E.;Ho, Yen-Yi

文献摘要

被引文献

相似文献

多变量计数数据在许多学科中很常见。这类数据中的变量通常表现出复杂的正或负依赖结构。通过同时考虑协变量依赖均值和相关性,我们提出了三种贝叶斯方法来建模双变量计数数据。直接方法利用了在Famoye(2010年,应用统计杂志)中开发的二元负二项概率质量函数。第二种方法使用双变量泊松-伽马混合模型间接拟合双变量计数数据。第三种方法是双变量高斯Copula模型。根据模拟分析的结果,间接和Copula方法在模型拟合和识别协变量依赖关联方面总体上比直接方法表现得更好。所提出的方法被应用于两个分别用于研究乳腺癌和黑色素瘤的RNA测序数据集(BRCA-US和SKCM-US),这些数据集是通过国际癌症基因组联盟获得的。
Multivariate count data are common in many disciplines. The variables in such data often exhibit complex positive or negative dependency structures. We propose three Bayesian approaches to modeling bivariate count data by simultaneously considering covariate-dependent means and correlation. A direct approach utilizes a bivariate negative binomial probability mass function developed in Famoye (2010,Journal of Applied Statistics). The second approach fits bivariate count data indirectly using a bivariate Poisson-gamma mixture model. The third approach is a bivariate Gaussian copula model. Based on the results from simulation analyses, the indirect and copula approaches perform better overall than the direct approach in terms of model fitting and identifying covariate-dependent association. The proposed approaches are applied to two RNA-sequencing data sets for studying breast cancer and melanoma (BRCA-US and SKCM-US), respectively, obtained through the International Cancer Genome Consortium.