Two-stage methodology for regression, path analysis, and structural equation models with item-level missingness
Two-stage methodology for regression, path analysis, and structural equation models with item-level missingness
批准号:
RGPIN-2015-05251
负责人:
Savalei, Victoria
金额:
$1.02万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31
中文摘要
拟议的研究的目标是开发和评估一种新的方法来处理不完整的数据,该方法适用于个别项目数据缺失的独特情况,但统计模型处于组合(项目总和)的水平。这种情况经常发生在社会科学领域。例如,在心理学中,模型中的变量通常是量表分数。如果变量“自尊”被用于回归,它将被计算为构成自尊量表的10个项目的总分。第二个应用是在结构方程模型(SEM)的背景下--复杂的多变量回归模型,它可能涉及潜在变量,并允许测试复杂的心理学理论。在这种情况下,相对于心理学中常用的典型样本量,模型往往较大。建议通过将项目组合成复合材料来减小模型尺寸。*拟议的研究将最近开发的用于不完全数据的两阶段(TS)方法(Savalei Savalei&Falk,2014)扩展到复合数据的情景。它被认为是一种不会因数据丢失而丢失任何信息的方法。上下文是连续的正态分布数据,扩展到非正态数据。缺失数据的现代处理方法包括最大似然(ML)和多重补偿(MI)。只要可用,ML就会更简单、更优雅,并提供独特的解决方案。然而,ML方法不适用于组合的情况,因为它不能处理没有直接进入模型的变量的缺失数据。两阶段(TS)过程是一种基于ML的过程,它将缺失数据处理与模型拟合分开。在这样做的时候,它允许直接处理复合材料。在阶段1中,应用最大似然估计来获得具有缺失数据的原始变量的均值和协方差矩阵的估计以及相关信息矩阵的估计。然后,将估计值转换为与复合材料的适当数量相对应。在阶段2中,模型适用于阶段1中的组合的均值和协方差矩阵,本质上将这些量视为来自完整数据。标准误差的一致估计是使用夹心估计器获得的,其中来自阶段1的复合材料的信息矩阵在夹心的“肉”中。*在这项研究中,我将把TS方法扩展到用于复合材料。我将与免费的R包的开发人员合作,在R中实现这种方法。因为这种方法即使在简单的回归中也很有用,所以我将编写一个R包,以允许在此上下文中不使用SEM包的简单实现。将对该方法进行广泛的评价。这一方法论的发展、研究和普及将极大地造福于从事不完全数据工作的社会科学家。*****
英文摘要
The proposed research has the goal of developing and evaluating a new methodology to deal with incomplete data, relevant for the unique situation when data are missing on the individual items, but the statistical model is at the level of composites (sums of items). This scenario often occurs in the social sciences. For example, in psychology the variables in the model are often scale scores. If the variable "self-esteem" is to be used in regression, it will be computed as the sum score of 10 items that comprise the self-esteem scale. A second application is in the context of structural equation models (SEMs)--complex multivariate regression models that may involve latent variables and allow for testing of complex psychological theories. In this case, models are often large relative to the typical sample sizes commonly used in psychology. It is recommended to reduce model size by combining items into composites. ***The proposed research extends the recently developed two-stage (TS) methodology for incomplete data (Savalei Savalei & Falk, 2014) to the scenario with composites. It is argued to be the method that does not lose any information due to missing data. The context is continuous normally distributed data, with extensions to nonnormal data. Modern approaches to missing data include maximum likelihood (ML) and multiple imputation (MI). Whenever available, ML is simpler and more elegant, and provides a unique solution. However, the ML methodology is not available for the case with composites as it cannot handle missing data on variables that do not directly enter the model. The two-stage (TS) procedure is an ML-based procedure that separates missing data treatment from model fitting. In doing so, it allows for a straight-forward treatment of composites. In Stage 1, ML is applied to obtain estimates of means and covariance matrix, as well as of the associated information matrix, of the original variables that have missing data. The estimates are then transformed to correspond to the appropriate quantities for the composites. In Stage 2, the model is fit to the means and covariance matrix of the composites from Stage 1, essentially treating these quantities as if they had come from complete data. Consistent estimates of standard errors are obtained using the sandwich estimator, where the information matrix for the composites from Stage 1 is in the "meat" of the sandwich. ***In this research, I will extend the TS methodology for use with composites. I will work with the developer of a free R package for SEM (lavaan) to implement this methodology in R. Because this methodology is useful even in simple regression, I will write an R package to allow a simple implementation in this context without the use of SEM packages. Extensive evaluations of the methodology will be conducted. The development, study, and popularization of this methodology will greatly benefit social scientists working with incomplete data. *****
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Improving Fit Assessment and Incomplete Data Diagnostics in Structural Equation Modeling
-
批准号:RGPIN-2021-02958
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2022
-
负责人:Savalei, Victoria
-
依托单位:
Improving Fit Assessment and Incomplete Data Diagnostics in Structural Equation Modeling
-
批准号:RGPIN-2021-02958
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2021
-
负责人:Savalei, Victoria
-
依托单位:
Two-stage methodology for regression, path analysis, and structural equation models with item-level missingness
-
批准号:RGPIN-2015-05251
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.02万
-
财政年份:2018
-
负责人:Savalei, Victoria
-
依托单位:
Two-stage methodology for regression, path analysis, and structural equation models with item-level missingness
-
批准号:RGPIN-2015-05251
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.02万
-
财政年份:2017
-
负责人:Savalei, Victoria
-
依托单位:
Two-stage methodology for regression, path analysis, and structural equation models with item-level missingness
-
批准号:RGPIN-2015-05251
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.02万
-
财政年份:2016
-
负责人:Savalei, Victoria
-
依托单位:
Two-stage methodology for regression, path analysis, and structural equation models with item-level missingness
-
批准号:RGPIN-2015-05251
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.02万
-
财政年份:2015
-
负责人:Savalei, Victoria
-
依托单位:
海外基金