Data Analysis Using Finite Mixture Models
Data Analysis Using Finite Mixture Models
批准号:
9404479
负责人:
Donald Rubin
金额:
$11.6万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1994
资助国家:
美国
项目状态:
已结题
起止时间:
1994-07-01 至 1998-06-30
中文摘要
有限混合模型在社会科学和医学科学中越来越流行,用于分析由分类类型组成的总体所产生的数据。本研究项目考虑使用有限混合模型分析数据的贝叶斯方法。将经典方法应用于此类模型是困难的,因为此类模型的概率是非标准的:由于混合成分的标记中的对称性,它们本质上是多峰的,即使对于单个固定的成分标记也可能是多峰的,并且不满足经典似然比检验所要求的正则性条件。感兴趣的主要科学问题涉及到关于模型参数的推断、确定混合组分的数量以及随机地将采样单元分类为混合组分。统计计算的现代进展,如EM和ECM算法、数据增强和Gibbs抽样,被用来获得关于模型参数的后验分布的推断;然而,由于上述困难,在应用这些方法时需要特别注意。本研究描述了区分上述两种模式的方法,以及在存在多个模式的情况下进行数据分析的方法。从后验分布中提取的模型参数可以用来获得从后验预测分布中提取的类似于当前实验的重复实验。测试统计量或其他差异度量的后验预测分布可用于评估模型的适合性,例如,即使问题是不规则的,也可以比较两组分和三组分混合模型。此外,对现有模型的合理备选方案的先验分布进行平均可用于估计评估现有模型相对于此类备选方案的适当性所需的样本量,因此可用于指导设计决策。考虑一类特殊的统计模型,称为混合模型,假设感兴趣的总体由许多相对同质的子总体组成,这种情况并不少见。当需要一个相对复杂的模型来描述在整个总体内观察到的数据模式时,混合模型是有用的,而相对简单的模型适用于每个子总体。经典方法在这种情况下的用处有限,例如,经典方法不适用于确定是否存在以及有多少子总体存在的关键问题。这项研究建议旨在开发使用混合模型分析数据和评估此类模型的充分性的新方法。这些新方法利用最新的理论和计算进展,对混合模型的重要特征做出准确的推断。基本方法是对数据所支持的所有人口描述进行平均,从而对人口的预期变化和模式作出准确的评估。
英文摘要
Finite mixture models are increasingly popular in the social and medical sciences for analyzing data thought to arise from a population consisting of categorical types. This research project considers Bayesian methods for analyzing data using finite mixture models. The application of classical methods to such models is difficult because the likelihoods for such models are nonstandard: they are inherently multimodal due to symmetries in the labeling of the mixture components, may be multimodal even for a single fixed labeling of the components, and fail to satisfy the regularity conditions required for classical likelihood ratio tests. The principal scientific questions of interest concern drawing inferences about model parameters, determining the number of mixture components, and stochastically classifying sampling units into the mixture components. Modern advances in statistical computing, such as the EM and ECM algorithms, data augmentation, and Gibbs sampling, are used to obtain inferences about the posterior distribution of model parameters; however special care is needed in applying these methods because of the difficulties mentioned above. This research describes methods for distinguishing between the two types of modes described above, and for carrying out the data analysis given the existence of multiple modes. Draws from the posterior distribution of the model parameters can be used to obtain draws from the posterior predictive distribution of replicate experiments similar to the current experiment. The posterior predictive distribution of test statistics, or of other discrepancy measures, can be used to evaluate the fit of a model, e.g., comparing two component and three component mixture models even though the problem is irregular. Additionally, averaging over a prior distribution on plausible alternatives to the existing model can be used to estimate the sample size required to assess the appropriateness of the existing model against such alternatives and therefore can be used to inform design decisions. It is not uncommon to consider a particular class of statistical models, called mixture models, that assume the population of interest consists of a number of relatively homogeneous subpopulations. Mixture models are useful when a relatively complex model would be required to describe the pattern of data that is observed within the entire population, whereas a relatively simple model applies within each subpopulation. Classical approaches are of limited use in such cases, e.g., classical methods do not apply to the crucial question of determining whether and how many subpopulations are in evidence. This research proposal aims to develop new methods for analyzing data using mixture models and for assessing the adequacy of such models. These new methods take advantage of recent theoretical and computational advances to draw accurate inferences about the important features of the mixture models. The basic approach is to average over all descriptions of the population that are supported by the data and thereby provide an accurate assessment of the variation and patterns to be expected in the population.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Generalized Propensity Score Methods
-
批准号:0550887
-
项目类别:Continuing Grant
-
资助金额:$13.51万
-
财政年份:2006
-
负责人:Donald Rubin
-
依托单位:
Multiple Imputation: Research for the Third Decade
-
批准号:9705158
-
项目类别:Continuing Grant
-
资助金额:$24.05万
-
财政年份:1997
-
负责人:Donald Rubin
-
依托单位:
Bridging Randomized Experiments and Observational Studies
-
批准号:9709359
-
项目类别:Continuing Grant
-
资助金额:$15.01万
-
财政年份:1997
-
负责人:Donald Rubin
-
依托单位:
Causal Inference Applied to Income Effects
-
批准号:9423018
-
项目类别:Standard Grant
-
资助金额:$11.08万
-
财政年份:1995
-
负责人:Donald Rubin
-
依托单位:
Applications of Modern Statistical Thinking to the Social Sciences
-
批准号:9207456
-
项目类别:Continuing Grant
-
资助金额:$27.27万
-
财政年份:1992
-
负责人:Donald Rubin
-
依托单位:
Mathematical Sciences: Topics in Nonparametric and Semiparametric Regression and Correlation Analysis
-
批准号:9106488
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:1991
-
负责人:Donald Rubin
-
依托单位:
Mathematical Sciences Research Equipment 1990
-
批准号:9005696
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:1990
-
负责人:Donald Rubin
-
依托单位:
Extending and Implementing Multiple Imputation Technology
-
批准号:8805433
-
项目类别:Continuing Grant
-
资助金额:$16.54万
-
财政年份:1988
-
负责人:Donald Rubin
-
依托单位:
Collaborative Research on the Recalibration of Categorical Data to Achieve Comparability Over Time
-
批准号:8311428
-
项目类别:Standard Grant
-
资助金额:$35.44万
-
财政年份:1983
-
负责人:Donald Rubin
-
依托单位:
A Consumer Health Policy Information and Resource Center For New York City
-
批准号:7923391
-
项目类别:Continuing Grant
-
资助金额:$20.77万
-
财政年份:1980
-
负责人:Donald Rubin
-
依托单位:
Planning a Technical Information Resource Center For Health Care Consumers
-
批准号:7821724
-
项目类别:Standard Grant
-
资助金额:$2.63万
-
财政年份:1978
-
负责人:Donald Rubin
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:USHARANI HAREESH GOVINDARA JAN
-
依托单位:
基于Meta-analysis的新疆棉花灌水增产模型研究
-
批准号:41601604
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2016
-
负责人:赵爱琴
-
依托单位:
大规模微阵列数据组的meta-analysis方法研究
-
批准号:31100958
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:赵洪雅
-
依托单位:
用“后合成核磁共振分析”(retrobiosynthetic NMR analysis)技术阐明青蒿素生物合成途径
-
批准号:30470153
-
项目类别:面上项目
-
资助金额:22.0万元
-
批准年份:2004
-
负责人:刘本叶
-
依托单位: