Valid Inference when Analytical Models are Approximations
Valid Inference when Analytical Models are Approximations
批准号:
1512084
负责人:
Linda Zhao
金额:
$53.2万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-07-15 至 2019-06-30
中文摘要
统计推理方法在现代生活中被用来回答问题。例如,是什么影响了一个城镇的犯罪率;哪些因素对房地产市场有重要影响;哪些基因与某种疾病有关;为了减缓气候变化,最需要控制的因素是什么?统计方法用于解决诸如此类的问题。然而,统计数据结构和为分析而开发的数学模型往往不一致。这个项目源于对标准推断分析和他们试图描述的世界统计数据之间不匹配的广泛的统计关注。本研究区分了传统上描述实验和观测数据中相关关系的统计模型和分析中使用的推理模型。为此,该项目研究了一种范式,在这种范式中,抽样模型意味着它们所描述的数据的真实结构的忠实表示。同时,应用于数据的分析模型只被视为对现实的近似描述。统计抽样表示不必与分析模型相匹配,尽管两者在某些重要方面应该协调一致。在准确的表征和忽略这一区别的经典程序所声称的之间存在着显著的差异。许多以前的统计研究人员已经注意到这种区别,并提出了各种部分适当的方法。然而,在研究的方向上澄清这种区别,然后追求结果会导致一种与通常用于关系和观测数据的推断理论有所不同的推断理论。承认并适当适应这种二元性,就会为一些重要的统计问题带来新的方法。一种这样的新方法是在随机临床试验的背景下,人们希望估计某些治疗相对于其他治疗或安慰剂对照的效果。另一种是在各种大数据环境中发生的半监督学习环境中。当前研究的核心是针对线性分析模型的设计。这包括对解释性协变量向量(x变量)和数值因变量(Y)的观察。解析模型将Y的最佳线性近似值作为X变量的线性函数。实际上没有对样本中的(X,Y)对做任何假设,除了它们从(X,Y)对的某些未知联合分布中形成一个统计样本并具有所需的低阶矩。“最佳”的概念是以统计上自然的方式定义的,与最小化预测误差的平方有关。由此可见,参数的普通最小二乘估计仍然具有理想的渐近性质。关于它们的(渐近)性能的推断可以通过标准的三明治估计量得到。然而,一个新导出的迭代对自举显示出更准确的实际样本量的推断信息。如果有更多关于X分布的信息(例如关于其均值和方差的知识),那么通常的最小二乘解可以得到改进。这一观察结果通过间接途径提出了改进随机临床试验中估计平均治疗效果的标准方法,以及在半监督学习环境中对数值结果进行线性预测的建议。在上述发展过程中,还暴露了各种其他问题。我们还计划研究上述设置的泛化-例如具有分类y变量(分类)的模型和其他广义线性分析模型。我们早期的研究涉及经典背景下的后选择推理,在这种背景下,数据模型及其分析是一致的,我们现在打算在当前的背景下追求类似的问题,在这种背景下,它们不一致。
英文摘要
Statistical inferential methods are used to answer questions throughout modern life. For example, what affects the crime rate in a town; which factors are important influences on the housing market; which genes are associated to a certain disease; what are the most important elements to control in order to mitigate climate change? Statistical methods are used to address questions such as these. However, often the statistical data structure and the mathematical model developed for the analysis do not agree. This project arises from a broadly based statistical concern about the mismatch between standard inferential analyses and the statistics of the world they are trying to describe. This research draws a distinction between the statistical models that conventionally describe the correlational relations in experimental and observational data and the inferential models that are used in their analysis. To this end, the project investigates a paradigm in which sampling models are meant to be faithful representations of the real-world structure of the data they are describing. At the same time, the analytical models to be applied to the data are viewed only as approximate descriptions of that reality. The statistical-sampling representations need not match the analytical models, though the two should harmonize in certain important respects. There is a significant disparity between accurate characterization and what is claimed by classical procedures that ignore this distinction. The distinction has been noted by many previous statistical researchers, and various partially adequate approaches have been suggested. Nevertheless, clarifying this distinction in the directions under study and then pursuing the consequences leads to a theory of inference somewhat different from that in common use for relational and observational data. Acknowledging and properly accommodating this duality then leads to new methodology for some important statistical problems. One such new methodology is within the setting of randomized clinical trials in which one wishes to estimate the effect of certain treatment(s) relative to others or to placebo controls. Another is within the setting of semi-supervised learning that occurs in various big-data contexts. The core of the current research is designed for linear analytical models. These involve observations on a vector of explanatory covariates (X-variables) and a numerical dependent variable (Y). The analytical model constructs the best linear approximant of Y as a linear function of the X variables. Virtually no assumptions are made about the (X,Y) pairs in the sample, other than that they form a statistical sample drawn from some unknown joint distribution of (X,Y) pairs and possess desired low-order moments. The notion of "best" is defined in a statistically natural fashion related to minimizing squared prediction error. It follows that the ordinary least squares estimators of parameters still have desirable asymptotic properties. Inference about their (asymptotic) performance can be derived via the standard sandwich estimator. However, a newly derived iterated pairs-bootstrap is shown to give substantially more accurate inferential information for realistic sample sizes. If more information is available about the distribution of X (such as knowledge of its mean and variance) then the usual least-squares solutions can be improved. This observation leads via an indirect path to suggestions that improve the standard methodology for estimating average treatment effect in randomized clinical trials and for producing linear predictions of numerical outcomes in settings of semi-supervised learning. Various additional issues are exposed in the course of the above developments. We also plan to investigate generalizations of the above setting -- for example to models having categorical Y-variables (classification) and to other generalized-linear analytical models. Our earlier research involved post-selection inference in the classical setting in which models for the data and its analysis coincide, and we now intend to pursue analogous issues in the current context in which they do not.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bayesian Inference Estimation in Nonparametric Regression and its Frequentist Properties
-
批准号:9971848
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:1999
-
负责人:Linda Zhao
-
依托单位:
海外基金