1 AN INFORMATION-GAIN MEASURE OF FIT IN PROC LOGISTIC

1 AN INFORMATION-GAIN MEASURE OF FIT IN PROC LOGISTIC
复制标题

1 过程物流中适合性的信息增益测量

DOI:
--
复制
发表时间:
1998
期刊:
--
影响因子:
--
通讯作者:
M. Barton
M. Barton
中科院分区:
--
文献类型:
--
作者:
E. Shtatland;M. Barton

文献摘要

被引文献

相似文献

本文是文献[1]1088页和文献[3],第413-414页)在1997年NESUG会议上发表的论文的续篇。在[1]中,建议使用([3],第832页)、Proc混合六([3],第547页,606信息增益或信息差异统计量作为607),等等。理解和适当解释这些统计数据的最自然和最有意义的标准所需的专门知识并不适合所有用户都能使用的这种流行的统计模型。因此,希望PROC REG、LOGISTICS、GENMOD等程序具有易于理解的PHREG和其他程序的通用标准。信息差统计的含义对于具有允许许多应用的非常方便的可加性的性质是常见的。SAS用户评估单个解释变量或一组解释变量在信息增益方面的贡献(以位为单位)。有趣的是,这些统计数据与检验统计学意义时使用的卡方统计数据直接相关。当我们对拟合标准感兴趣时,这一点尤其重要:R-平方和基于理解统计似然之间的关系的统计量。意义和实质意义或重要性。在[1]中,我们主要关注PROC REG。在这里,最受关注的是卫生保健研究中最常用的回归方法--Proc Logistic案例。下面给出的大部分材料都与PROC REG和PROC LOGISTICS同等相关。此外,我们将发现这两种类型的回归分析在统计意义和实质重要性的相互作用方面有惊人的相似之处,如果两者都使用相同的“货币”信息来衡量的话。。SAS用户和SAS统计建模过程使用SAS统计建模过程的最困难的点之一(在选择模型类型之后)是理解我们所建立的模型在某种拟合优度标准方面是否足够好。这意味着,至少,我们必须理解和解释打印输出。这项任务的难度是复杂的,因为,例如,PROC REG至少有11个关于FIT的统计数据([2],第1369页),PROC GENMOD 9个统计数据([3],第273页),PROC LOGISTICS 4([2],可能担任通用适配标准的候选人)。
This paper is a continuation of the paper [1] presented p. 1088, and [3], pp. 413-414), PROC PHREG three at NESUG, 1997. In [1] it was proposed to use ([3], p. 832), PROC MIXED six ([3], pp. 547, 606information-gain or information-difference statistics as 607), and so on. The expertise required to nderstand the most natural and most meaningful criteria of and interpret properly each of these statistics is not goodness of fit in such popular statistical modeling available to all users. Therefore it would be desirable to procedures as PROC REG, LOGISTIC, GENMOD, have a universal criterion with an easy to understand PHREG and others. The information-difference statistics meaning that would be common for have a very convenient property of additivity that allows many applications. SAS users to evaluate the contribution of an individual explanatory variable or a group of explanatory variables in terms of information gain in bits. It is interesting that these statistics are directly related to For Chi-Square statistics used in testing statistical significance. This is especially important when we are interested in criterion of fit: R-Square and a statistic based on understanding the relationship between statistical likelihood. significance and substantive significance or importance. In [1] we focused mostly on PROC REG. Here the case of PROC LOGISTIC, the most popular regression procedure in health care research, is paid more attention. Much of the material given below is equally related to both PROC REG and PROC LOGISTIC. Moreover, we will find a striking analogy between these two types of regression analysis regarding the interplay of statistical significance and substantive importance if both are measured by using the same ‘currency’ information. . SAS USERS AND SAS STATISTICAL MODELING PROCEDURES One of the most difficult points in using SAS statistical modeling procedures (after the type of the model is chosen) is to understand whether the model we have built is good enough in terms of some goodness-of-fit criterion. This means that at a minimum, we have to understand and interpret printouts. The difficulty of this task is compounded because, for example, PROC REG has at least eleven statistics of fit ([2], p. 1369), PROC GENMOD nine statistics ([3], p. 273), PROC LOGISTIC four ([2], POSSIBLE CANDIDATES FOR THE ROLE OF A UNIVERSAL CRITERION OF FIT There are two clear candidates for the role of a universal