课题基金 / 基金详情

Mathematical Methods for Small Sample Biostatistical Inference

Mathematical Methods for Small Sample Biostatistical Inference
小样本生物统计推断的数学方法
批准号:
0092659
负责人:
John Kolassa
金额:
$12.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2000
资助国家:
美国
项目状态:
已结题
起止时间:
2000-09-01 至 2004-08-31

项目摘要

项目成果

John Kolassa的其他基金

相似基金

相关文献

中文摘要
翻译
对由离散分布产生的指数族进行条件推理的方法多种多样。正态理论方法依赖于数据集中汇总统计量联合分布的近似多元正态性,对于小数据集通常是不准确的,对于表明大参数影响的汇总,它们的质量通常很差。他们也忽略了数据的离散性。更复杂的近似技术,被称为鞍点技术,经常在正常理论方法不充分的情况下使用。这些技术通常没有考虑到数据的离散性,因此在未经修改的形式下是次优的。精确的推断技术也是可用的,但是这些技术只适用于有限数量的模型,需要专有软件,并且当样本量达到中等大小时失败。该软件的扩展,采用蒙特卡罗技术为更大的样本量还没有商用。这些蒙特卡罗技术还有另一个缺点,即为同一数据集提供各种结果。所提出的技术使用鞍点近似,以一种解释数据离散性的方式,同时避免了精确计算中大多数计算棘手的方面。在这项拨款申请中提出的一些项目涉及新的近似,例如近似高维分布函数,而其他项目涉及对现有近似的修改,以避免数值不稳定性。其他项目包括制定置信区域以使准确校准变得容易,修改条件事件以获得更强大的分析,并执行诊断以确保使用正确的近似值。这些方法具有足够的通用性,可以应用于晶格上支持的任何正则指数族,因此也可以应用于任何具有正则链接的广义线性模型,晶格上支持的观测值,以及条目限制在晶格内的设计矩阵。将适用的模型的例子是逻辑回归,泊松回归,包括列联表的对数线性模型和多项模型。回归模型与更奇特的误差结构,包括正泊松和负二项分布,也将适应。本研究旨在帮助对多个参数的统计推断,在存在其他不直接感兴趣的干扰参数的情况下,当分布模型是离散的。例如,癌症患者病情持续缓解的概率可以建模为多种因素的函数。其中一些影响,比如病人接受了哪种治疗,或者病人是否有其他癌症相关的病理,可能会推广到其他人群,而另一些影响,比如病人接受治疗的特定中心的影响,可能不会推广到其他人群。因此,人们可能对描述感兴趣的参数所取的可能值感兴趣,而不需要同时估计其余参数。通常,人们将与麻烦参数相关的信息视为固定不变的,并根据这些信息有条件地执行推理。也就是说,通过将实验结果与可能结果的总体进行比较,从而使有关有害参数的信息保持固定,从而评估有关感兴趣参数的证据。本文提出的研究议程提出了进行这些计算的方法,这些方法平衡了精确方法的高计算成本和近似值的潜在不准确性,并引入和结合了精确和近似计算的新方法。这些新方法将使应用科学中经常出现的小型离散数据集的分析更快、更准确。
英文摘要
A wide variety of techniques exist for conditional inference on exponential families arising from discrete distributions. Normal theory methods, which rely on the approximate multivariate normality of the joint distribution of summary statistics from the data set, are often inaccurate for small data sets, and their quality can often be poor for summaries that indicate large parameter effects. They also ignore discreteness in the data. More sophisticated approximation techniques, known as saddlepoint techniques, are often used in cases when normal theory methods are inadequate. These techniques often do not account for discreteness in data, and hence are suboptimal in their unmodified forms. Exact inferential techniques are also available, but these techniques apply only to a limited number of models, require proprietary software, and fail when sample size reaches a moderate size. Extensions to this software that employ Monte Carlo techniques for larger sample sizes are not yet commercially available. These Monte Carlo techniques have the further disadvantage of delivering a variety of results for the same data set. The techniques proposed use saddlepoint approximations in a way that accounts for discreteness in the data while avoiding most of the computationally intractable aspects of exact calculations. Some of the projects proposed in this grant application involve new approximations, such as for approximating higher--dimensional distribution functions, and others involve modifications to existing approximations to avoid numerical instabilities. Other projects involve formulating confidence regions to make accurate calibration easy, and modifying the conditioning event to obtain a more powerful analysis, and performing diagnostics to ensure that the proper approximations are used. These methods will be general enough to apply to any canonical exponential family supported on a lattice, and hence to any generalized linear model with canonical link, observations supported on a lattice, and design matrix whose entries are confined to a lattice. Examples of models that will be accommodated are logistic regression, Poisson regression including log linear models for contingency tables, and multinomial models. Regression models with more exotic error structures, including positive Poisson and negative binomial distributions, will also be accommodated.This proposed research is intended to aid in statistical inference on multiple parameters, in the presence of other nuisance parameters that are not of direct interest, when the distribution modeled is discrete. For example, the probability that a cancer patient will stay in remission can be modeled as a function of a variety of factors. Some of these effects, like which treatment a patient received or whether the patient had other cancer--related pathologies, may generalize to other populations, and others, like the effect of a particular center where the patient was treated, may not generalize. Thus one might be interested in describing the possible values that the parameters of interested take on, without being required to simultaneously estimate the remaining parameters. Typically one treats information associated with nuisance parameters as held fixed, and performs inference conditionally on this information. That is, one assesses the the evidence concerning the parameter of interest by comparing experimental results to the population of possible results such that the information about nuisance parameters is held fixed. The research agenda proposed here presents methods for doing these calculations, which balance high computational costs of exact methods against potential inaccuracies of approximations, and introduces and combines new methods for bothexact and approximate calculations. These new methods will make the analysis of small discrete data sets, commonly occurring in applied sciences, quicker and more accurate.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Higher-Order Asymptotics and Accurate Inference for Post-Selection
  • 批准号:
    1712839
  • 项目类别:
    Standard Grant
  • 资助金额:
    $12.0万
  • 财政年份:
    2017
  • 负责人:
    John Kolassa
  • 依托单位:
Mathematical Methods for Approximately Exact Statistical Inference
  • 批准号:
    0906569
  • 项目类别:
    Standard Grant
  • 资助金额:
    $12.26万
  • 财政年份:
    2009
  • 负责人:
    John Kolassa
  • 依托单位:
Mathematical Methods for Small--Sample Biostatistical Inference
  • 批准号:
    0505499
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2005
  • 负责人:
    John Kolassa
  • 依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data