Multivariate mixtures of Polya trees for modelling ROC data

Multivariate mixtures of Polya trees for modelling ROC data
复制标题

DOI:
10.1177/1471082x0700800106
复制
发表时间:
2008-04-01
影响因子:
1
通讯作者:
Gardner, Ian A.
Gardner, Ian A.
中科院分区:
数学4区
文献类型:
--
作者:
Hanson, Timothy E.;Branscum, Adam J.;Gardner, Ian A.

文献摘要

被引文献

相似文献

受试者工作特征(ROC)曲线提供了诊断测试准确性的图形度量。由于ROC曲线是使用未感染和感染人群的诊断测试结果的分布来确定的,因此为这些成分分布开发灵活模型的趋势越来越大。我们提出了从多变量血清学数据对几个ROC曲线进行联合非参数估计的方法。我们开发了一种经验贝叶斯方法,该方法允许使用有限Polya树先验的贝叶斯多元混合建模任意非感染和感染组件分布。获得了健壮的、数据驱动的ROC曲线和曲线下面积推断,并提出了一种简单的方法来测试Dirichlet过程与更一般的Polya树模型。当使用Polya树对显示聚类的大型多变量数据集建模时,可能会出现计算挑战。我们讨论并实施了解决这些障碍的实际程序,这些程序应用于用于评估两种ELISA检测约翰氏病的性能的双变量数据。
Receiver operating characteristic (ROC) curves provide a graphical measure of diagnostic test accuracy. Because ROC curves are determined using the distributions of diagnostic test outcomes for noninfected and infected populations, there is an increasing trend to develop flexible models for these component distributions. We present methodology for joint nonparametric estimation of several ROC curves from multivariate serologic data. We develop an empirical Bayes approach that allows for arbitrary noninfected and infected component distributions that are modelled using Bayesian multivariate mixtures of finite Polya trees priors. Robust, data-driven inferences for ROC curves and the area under the curve are obtained, and a straightforward method for testing a Dirichlet process versus a more general Polya tree model is presented. Computational challenges can arise when using Polya trees to model large multivariate data sets that exhibit clustering. We discuss and implement practical procedures for addressing these obstacles, which are applied to bivariate data used to evaluate the performances of two ELISA tests for detection of Johne's disease.