课题基金 / 基金详情

Valid Inference when Analytical Models are Approximations

Valid Inference when Analytical Models are Approximations
当分析模型为近似值时的有效推理
批准号:
1512084
负责人:
Linda Zhao
金额:
$53.2万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-07-15 至 2019-06-30

项目摘要

项目成果

Linda Zhao的其他基金

相似基金

相关文献

中文摘要
翻译
统计推断方法被用来回答整个现代生活中的问题。 例如,是什么影响了一个城镇的犯罪率;哪些因素对房地产市场有重要影响;哪些基因与某种疾病有关;为了减缓气候变化,最重要的控制因素是什么? 统计方法被用来解决诸如此类的问题。然而,统计数据结构和为分析而开发的数学模型往往不一致。这个项目源于一个广泛的统计问题,即标准推理分析与他们试图描述的世界统计数据之间的不匹配。本研究区分了传统上描述实验和观测数据中相关关系的统计模型和用于分析的推理模型。为此,该项目研究了一种范式,在这种范式中,采样模型意味着它们所描述的数据的真实世界结构的忠实表示。 与此同时,将应用于数据的分析模型仅被视为对这一现实的近似描述。 尽管在某些重要的方面两者应该协调一致,但纵向抽样的表述方式不需要与分析模型相匹配。有一个显着的差距之间的准确表征和经典的程序,忽略了这一区别声称。许多以前的统计研究人员已经注意到了这种区别,并提出了各种部分适当的方法。然而,在研究的方向上澄清这种区别,然后追求结果,导致一种推理理论,与通常用于关系和观察数据的理论有些不同。承认并适当地适应这种二重性,然后导致一些重要的统计问题的新方法。一种这样的新方法是在随机临床试验的背景下,其中人们希望估计某些治疗相对于其他治疗或安慰剂对照的效果。另一种是在各种大数据背景下发生的半监督学习的设置中。当前研究的核心是线性分析模型。这些涉及对解释协变量(X变量)和数值因变量(Y)向量的观察。分析模型将Y的最佳线性逼近构造为X变量的线性函数。实际上,对样本中的(X,Y)对没有任何假设,除了它们形成从(X,Y)对的某些未知联合分布中抽取的统计样本并具有期望的低阶矩。“最佳”的概念是以统计上自然的方式定义的,与最小化平方预测误差有关。由此可见,参数的普通最小二乘估计仍然具有令人满意的渐近性质。关于它们的(渐近)性能的推断可以通过标准三明治估计量来推导。然而,一个新衍生的迭代配对引导显示,实际样本量,以提供更准确的推理信息。如果有更多关于X分布的信息(例如其均值和方差的知识),那么通常的最小二乘解可以得到改进。这一观察结果通过间接途径提出了建议,这些建议改进了在随机临床试验中估计平均治疗效果的标准方法,并在半监督学习环境中产生数值结果的线性预测。在上述发展过程中还暴露出各种其他问题。我们还计划研究上述设置的推广-例如具有分类Y变量(分类)的模型和其他广义线性分析模型。我们早期的研究涉及经典环境中的后选择推理,在经典环境中,数据及其分析的模型是一致的,我们现在打算在当前的背景下追求类似的问题,而它们并不一致。
英文摘要
Statistical inferential methods are used to answer questions throughout modern life. For example, what affects the crime rate in a town; which factors are important influences on the housing market; which genes are associated to a certain disease; what are the most important elements to control in order to mitigate climate change? Statistical methods are used to address questions such as these. However, often the statistical data structure and the mathematical model developed for the analysis do not agree. This project arises from a broadly based statistical concern about the mismatch between standard inferential analyses and the statistics of the world they are trying to describe. This research draws a distinction between the statistical models that conventionally describe the correlational relations in experimental and observational data and the inferential models that are used in their analysis. To this end, the project investigates a paradigm in which sampling models are meant to be faithful representations of the real-world structure of the data they are describing. At the same time, the analytical models to be applied to the data are viewed only as approximate descriptions of that reality. The statistical-sampling representations need not match the analytical models, though the two should harmonize in certain important respects. There is a significant disparity between accurate characterization and what is claimed by classical procedures that ignore this distinction. The distinction has been noted by many previous statistical researchers, and various partially adequate approaches have been suggested. Nevertheless, clarifying this distinction in the directions under study and then pursuing the consequences leads to a theory of inference somewhat different from that in common use for relational and observational data. Acknowledging and properly accommodating this duality then leads to new methodology for some important statistical problems. One such new methodology is within the setting of randomized clinical trials in which one wishes to estimate the effect of certain treatment(s) relative to others or to placebo controls. Another is within the setting of semi-supervised learning that occurs in various big-data contexts. The core of the current research is designed for linear analytical models. These involve observations on a vector of explanatory covariates (X-variables) and a numerical dependent variable (Y). The analytical model constructs the best linear approximant of Y as a linear function of the X variables. Virtually no assumptions are made about the (X,Y) pairs in the sample, other than that they form a statistical sample drawn from some unknown joint distribution of (X,Y) pairs and possess desired low-order moments. The notion of "best" is defined in a statistically natural fashion related to minimizing squared prediction error. It follows that the ordinary least squares estimators of parameters still have desirable asymptotic properties. Inference about their (asymptotic) performance can be derived via the standard sandwich estimator. However, a newly derived iterated pairs-bootstrap is shown to give substantially more accurate inferential information for realistic sample sizes. If more information is available about the distribution of X (such as knowledge of its mean and variance) then the usual least-squares solutions can be improved. This observation leads via an indirect path to suggestions that improve the standard methodology for estimating average treatment effect in randomized clinical trials and for producing linear predictions of numerical outcomes in settings of semi-supervised learning. Various additional issues are exposed in the course of the above developments. We also plan to investigate generalizations of the above setting -- for example to models having categorical Y-variables (classification) and to other generalized-linear analytical models. Our earlier research involved post-selection inference in the classical setting in which models for the data and its analysis coincide, and we now intend to pursue analogous issues in the current context in which they do not.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bayesian Inference Estimation in Nonparametric Regression and its Frequentist Properties
  • 批准号:
    9971848
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    1999
  • 负责人:
    Linda Zhao
  • 依托单位:
海外基金