Addressing data sparseness in contextual population research - Using cluster analysis to create synthetic neighborhoods

Addressing data sparseness in contextual population research - Using cluster analysis to create synthetic neighborhoods
复制标题

DOI:
10.1177/0049124106292362
复制
发表时间:
2007-02-01
影响因子:
6.3
通讯作者:
Wheaton, Blair
Wheaton, Blair
中科院分区:
法学2区
文献类型:
--
作者:
Clarke, Philippa;Wheaton, Blair

文献摘要

被引文献

相似文献

基于人口的调查数据的多级建模的使用通常受到每级2单元的少量病例的限制,这促使邻域文献中最近出现了应用聚类技术来解决数据稀疏问题的趋势。在这项研究中,作者使用蒙特卡罗模拟来研究边际组大小对多层模型性能、偏差和效率的影响。然后,他们使用聚类分析技术来最小化数据稀疏性,并检查模拟中的结果。他们发现,在数据稀疏的极端情况下,固定效应的估计是稳健的,而聚类分析是增加群体规模和防止方差成分高估的有效策略。然而,由于引入了人为的组内异质性,研究人员应该谨慎使用这种聚类技术的程度。
The use of multilevel modeling with data from population-based surveys is often limited by the small number of cases per Level 2 unit, prompting a recent trend in the neighborhood literature to apply cluster techniques to address the problem of data sparseness. In this study, the authors use Monte Carlo simulations to investigate the effects of marginal group sizes on multilevel model performance, bias, and efficiency. They then employ cluster analysis techniques to minimize data sparseness and examine the consequences in the simulations. They find that estimates of the fixed effects are robust at the extremes of data sparseness, while cluster analysis is an effective strategy to increase group size and prevent the overestimation of variance components. However, researchers should be cautious about the degree to which they use such clustering techniques due to the introduction of artificial within-group heterogeneity.