1 AN INFORMATION-GAIN MEASURE OF FIT IN PROC LOGISTIC
1 AN INFORMATION-GAIN MEASURE OF FIT IN PROC LOGISTIC
复制标题
1 过程物流中适合性的信息增益测量
DOI:
--
复制
发表时间:
1998
期刊:
影响因子:
--
通讯作者:
M. Barton
中科院分区:
文献类型:
--
作者:
E. Shtatland;M. Barton
This paper is a continuation of the paper [1] presented p. 1088, and [3], pp. 413-414), PROC PHREG three at NESUG, 1997. In [1] it was proposed to use ([3], p. 832), PROC MIXED six ([3], pp. 547, 606information-gain or information-difference statistics as 607), and so on. The expertise required to nderstand the most natural and most meaningful criteria of and interpret properly each of these statistics is not goodness of fit in such popular statistical modeling available to all users. Therefore it would be desirable to procedures as PROC REG, LOGISTIC, GENMOD, have a universal criterion with an easy to understand PHREG and others. The information-difference statistics meaning that would be common for have a very convenient property of additivity that allows many applications. SAS users to evaluate the contribution of an individual explanatory variable or a group of explanatory variables in terms of information gain in bits. It is interesting that these statistics are directly related to For Chi-Square statistics used in testing statistical significance. This is especially important when we are interested in criterion of fit: R-Square and a statistic based on understanding the relationship between statistical likelihood. significance and substantive significance or importance. In [1] we focused mostly on PROC REG. Here the case of PROC LOGISTIC, the most popular regression procedure in health care research, is paid more attention. Much of the material given below is equally related to both PROC REG and PROC LOGISTIC. Moreover, we will find a striking analogy between these two types of regression analysis regarding the interplay of statistical significance and substantive importance if both are measured by using the same ‘currency’ information. . SAS USERS AND SAS STATISTICAL MODELING PROCEDURES One of the most difficult points in using SAS statistical modeling procedures (after the type of the model is chosen) is to understand whether the model we have built is good enough in terms of some goodness-of-fit criterion. This means that at a minimum, we have to understand and interpret printouts. The difficulty of this task is compounded because, for example, PROC REG has at least eleven statistics of fit ([2], p. 1369), PROC GENMOD nine statistics ([3], p. 273), PROC LOGISTIC four ([2], POSSIBLE CANDIDATES FOR THE ROLE OF A UNIVERSAL CRITERION OF FIT There are two clear candidates for the role of a universal