ROC Curve Methodology
ROC Curve Methodology
批准号:
7594226
负责人:
Enrique Schisterman
金额:
$14.09万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AccountingAcute DiseaseAmericanAreaAttenuatedBiological AssayBiological MarkersBiomedical ResearchBiotechnologyChronicComputer softwareConfidence IntervalsDataDevelopmentDevicesDiagnosisDiagnosticDiagnostic testsDiseaseDisease modelEarly DiagnosisEffectivenessEvaluationIndividualInfertilityJointsJournalsLaboratoriesLeadLiteratureMeasurementMeasuresMethodologyMethodsModelingNormal Statistical DistributionNumbersOutcomeOxidative StressPatientsPeer ReviewPerformancePopulationPreventionProbabilityProcessPropertyPublicationsROC CurveRangeReactionReceiver Operating CharacteristicsResearchResearch PersonnelSamplingScreening procedureSelection BiasSourceSpecificitySpecimenStandards of Weights and MeasuresStatistical MethodsTechniquesTestingWorkbasecostdesigndesirediagnostic accuracyimprovedindexinginterestnovel diagnosticsrapid growthsimulationstatisticsthiobarbituric acidtool
中文摘要
正如有许多氧化应激的标志物一样,生物技术的快速发展意味着研究人员必须越来越多地考虑在他们的研究中使用哪种筛查或诊断测试。我使用ROC曲线的工作旨在为做出这些选择提供基于证据的方法。ROC曲线同时绘制了在不同测试截止点正确诊断的异常和正常受试者的比例。这种图形显示有助于选择最佳阈值,并能够轻松比较不同测试的能力。ROC曲线越来越多地用于基于群体的设置,而不是在某种程度上预先筛选个体的设置。然而,ROC曲线方法并没有被开发来解释常见的问题,例如缺失数据、测量误差、线性组合、混杂、参考偏倚、LOD和其他挑战。
我们已经提出了K样本U统计量(其中ROC曲线下面积(AUC)是一种特殊情况)的平均值的估计,当在某些采样单元中缺少感兴趣的结果的数据时,辅助变量在整个样本中可用。建议的估计利用辅助设备中的信息,而不需要假设的辅助设备和结果的联合分布。所提出的估计量的性质来自于一般结果的有效半参数估计的平均值的K样本U-统计量与随机结果,观察到的辅助变量和已知的missingness概率。
随机测量误差可以削弱生物标志物区分患病和非患病群体的能力。我们提出了一种方法,用于估计约登指数,AUC及其相关的最佳临界点的正态分布的生物标志物,校正正态分布的随机测量误差。我们还开发了这些校正估计的置信区间,通过模拟各种情况,使用三角洲方法和覆盖概率。将这些技术应用于生物标志物硫代巴比妥酸反应物质(TBARS),一种氧化应激的测量方法,已被提议作为不孕症的鉴别测量方法,在最佳临界点处的诊断有效性增加了50%。这一结果可能会导致生物标志物,曾经天真地认为无效成为有用的诊断设备。
由于多个标记物通常可用,我们考虑将它们组合以提高诊断准确性。由Su和Liu(1993)推导出的使AUC最大化的线性组合在特定的期望特异性范围内可能具有不令人满意的低灵敏度。我们考虑了在特异性范围内灵敏度的最大化,并提出了在高(或低)特异性范围内具有更高灵敏度的替代线性组合。此外,我们评估了协变量对该线性组合的影响,假设多个标志物或其转换遵循多变量正态分布。我们估计了这种标记物线性组合的ROC曲线,该曲线针对协变量和相应AUC的近似置信区间进行了调整。
在评估新诊断测试的研究中经常遇到的另一个问题是,由于测试的费用和/或侵入性,并非所有患者都接受疾病验证。事实上,让患者接受验证检测的决定通常取决于新检测的结果和疾病状态的其他预测因素。对于AUC估计仅基于具有经验证的疾病状态的患者的诊断测试,通常的估计值存在偏倚。我们开发了调整这种偏差的估计器。
当疾病状态信息缺失时,有必要对缺失数据或导致缺失的过程进行建模,以获得表现良好的AUC估计值。我们已经描述了一个双稳健估计,当疾病或缺失模型正确时,该估计是无偏的。该估计器不需要EM型迭代,并且易于使用标准软件计算。 它可以容纳离散和连续的标记,并允许选择验证的可能性是不可重复的。此外,双鲁棒估计提供了更多的保护,对模型误设定比其他目前可用的方法。
我们已经应用了上述方法,以表明TBARS具有超越机会的辨别能力。这项工作已经在同行评审的期刊上发表了23篇论文,包括《生物统计学》和《美国统计协会杂志》。
英文摘要
Just as there are many markers of oxidative stress, the rapid growth of biotechnology means that researchers increasingly must consider which screening or diagnostic test to use in their research. My work with ROC curves is aimed at providing evidence-based approaches for making these choices. The ROC curve simultaneously plots the proportion of both abnormal and normal subjects correctly diagnosed at various test cutoff points. This graphical display facilitates the selection of an optimal threshold and enables easy comparison of the abilities of different tests. Increasingly, ROC curves are used in population based settings as opposed to settings where individuals have been pre-screened to some degree. However, ROC curve methods were not developed to account for common problems such as missing data, measurement error, linear combinations, confounding, referral bias, LODs, and other challenges.
We have proposed estimators of the mean of a K-sample U-statistic (of which the area under the ROC curve (AUC) is a special case) when data on the outcomes of interest are missing in some sampled units and auxiliary variables are available in the entire sample. The proposed estimators exploit the information available in the auxiliaries without requiring assumptions about the joint distribution of the auxiliaries and outcomes. The properties of the proposed estimators are derived from general results on efficient semi-parametric estimation of the mean of a K-sample U-statistic with missing at random outcomes, observed auxiliary variables and known missingness probabilities.
Random measurement error can attenuate a biomarkers ability to discriminate between diseased and non-diseased populations. We present an approach for estimating the Youden index, the AUC and its associated optimal cut-point for a normally distributed biomarker that corrects for normally distributed random measurement error. We also developed confidence intervals for these corrected estimates using the delta method and coverage probability through simulation of a variety of situations. Applying these techniques to the biomarker thiobarbituric acid reaction substance (TBARS), a measure of oxidative stress that has been proposed as a discriminating measurement for infertility, yields a 50% increase in diagnostic effectiveness at the optimal cut-point. This result may lead to biomarkers that were once naively considered ineffective becoming useful diagnostic devices.
Since multiple markers are often available, we considered combining them to improve diagnostic accuracy. The linear combinations derived by Su and Liu (1993) that maximize the AUC may have unsatisfactorily low sensitivity over a certain range of desired specificity. We considered maximization of sensitivity over a range of specificity, and presented alternative linear combinations that have higher sensitivity over a range of high (or low) specificity. Additionally, we evaluated covariate effects on this linear combination assuming that the multiple markers or a transformation thereof, follow a multivariate normal distribution. We estimated the ROC curve of this linear combination of markers adjusted for covariates and approximate confidence intervals for the corresponding AUC.
Another frequently encountered problem in studies that evaluate new diagnostic tests is that not all patients undergo disease verification due to the expense and/or invasiveness of the test. In fact, the decision to subject patients to verification testing often depends on the results of the new test and other predictors of disease status. For diagnostic tests where AUC estimation is based only on patients with verified disease status, the usual estimators are biased. We developed estimators that adjust for this bias.
When information on disease status is missing, it is necessary either to model the missing data or the process leading to the missingness to obtain well-behaved estimators of the AUC. We have described a doubly robust estimator that is unbiased when the model for disease or the missingness is correct. This estimator does not require EM-type iterations and is easy to compute using standard software. It can accommodate both discrete and continuous markers and allows for the possibility that selection to verification is non-ignorable. In addition, the doubly robust estimator offers more protection against model misspecification than other currently available methods.
We have applied the methods described above to show that TBARS, has discriminating abilities above and beyond chance. This work has yielded 23 publications in peer reviewed journals including Biometrika and the Journal of the American Statistical Association.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ROC Curve Methodology
-
批准号:7734777
-
项目类别:
-
资助金额:$3.88万
-
财政年份:--
-
负责人:Enrique Schisterman
-
依托单位:
EAGeR Trial - The Effects of Aspirin in Gestation and Reproduction Trial
-
批准号:7734799
-
项目类别:
-
资助金额:$45.23万
-
财政年份:--
-
负责人:Enrique Schisterman
-
依托单位:
EAGeR Trial - The Effects of Aspirin in Gestation and Reproduction Trial
-
批准号:7594250
-
项目类别:
-
资助金额:$14.38万
-
财政年份:--
-
负责人:Enrique Schisterman
-
依托单位:
海外基金