The importance of distribution-choice in modeling substance use data: a comparison of negative binomial, beta binomial, and zero-inflated distributions.

The importance of distribution-choice in modeling substance use data: a comparison of negative binomial, beta binomial, and zero-inflated distributions.
复制标题

DOI:
10.3109/00952990.2015.1056447
复制
发表时间:
2015
期刊:
The American journal of drug and alcohol abuse
影响因子:
--
通讯作者:
Mikulich-Gilbertson S
Mikulich-Gilbertson S
中科院分区:
其他
文献类型:
--
作者:
Wagner B;Riggs P;Mikulich-Gilbertson S

文献摘要

相似文献

重要的是要正确理解多种药物成瘾之间的关联,以及同时发生的物质使用和精神疾病之间的关联。具体物质的结果(例如使用大麻的天数)具有分布特征,其范围很广,取决于所评价的物质和样本。我们推荐一个由四部分组成的策略来确定建模物质使用数据的适当分布。我们证明了这一战略,通过比较模型拟合和推论,从应用四种不同的分布模型使用的物质,范围很大,在其使用的流行率和频率。使用来自先前发表的研究的时间轴跟踪(TLFB)数据,我们使用负二项分布、β-二项分布及其零膨胀对应物来模拟大麻、香烟、酒精和阿片类药物使用治疗期间的天数比例。使用统计模型选择标准、视觉图和所得推断的比较评价了每个分布的拟合。我们证明了每种物质单独建模的可行性和实用性,并表明没有一个单一的分布提供了所有物质的最佳拟合。关于每种物质的使用和与重要临床变量的关联的推断在不同模型中不一致,并且因物质而异。因此,必须仔细选择和评估用于建模物质使用的分布,因为它可能会影响得出的结论。此外,汇总不同物质使用情况的通用程序可能并不理想。
It is important to correctly understand the associations among addiction to multiple drugs and between co-occurring substance use and psychiatric disorders. Substance-specific outcomes (e.g. number of days used cannabis) have distributional characteristics which range widely depending on the substance and the sample being evaluated. We recommend a four-part strategy for determining the appropriate distribution for modeling substance use data. We demonstrate this strategy by comparing the model fit and resulting inferences from applying four different distributions to model use of substances that range greatly in the prevalence and frequency of their use. Using timeline followback (TLFB) data from a previously-published study, we used negative binomial, beta-binomial and their zero-inflated counterparts to model proportion of days during treatment of cannabis, cigarettes, alcohol, and opioid use. The fit for each distribution was evaluated with statistical model selection criteria, visual plots and a comparison of the resulting inferences. We demonstrate the feasibility and utility of modeling each substance individually and show that no single distribution provides the best fit for all substances. Inferences regarding use of each substance and associations with important clinical variables were not consistent across models and differed by substance. Thus, the distribution chosen for modeling substance use must be carefully selected and evaluated because it may impact the resulting conclusions. Furthermore, the common procedure of aggregating use across different substances may not be ideal.