Model based clustering of large data sets: Tracing the development of spelling ability

Model based clustering of large data sets: Tracing the development of spelling ability
复制标题

DOI:
10.1007/bf02295648
复制
发表时间:
2004-09-01
期刊:
影响因子:
3
通讯作者:
Notenboom, A
Notenboom, A
中科院分区:
心理学4区
文献类型:
--
作者:
Hoijtink, H;Notenboom, A

文献摘要

被引文献

相似文献

关于拼写能力的发展,主要有两种理论:阶段模型和重叠波模型。本文将采用探索性模型为基础的聚类分析超过3500名小学生的反应子集的245个项目。为了评估这两种理论,将使用外部标准沿着发展维度对所产生的集群进行排序。给出了三个统计问题的解决方案:(1)一个能够处理大数据集并且只呈现非退化聚类的算法;(2)一个拟合优度检验,它不受可能响应向量的数量远远超过观测响应向量的数量这一事实的影响;以及(3)一种新的技术,数据消去,如果缺失数据机制是已知的,则可用于评估拟合优度检验。
There are two main theories with respect to the development of spelling ability: the stage model and the model of overlapping waves. In this paper exploratory model based clustering will be used to analyze the responses of more than 3500 pupils to subsets of 245 items. To evaluate the two theories, the resulting clusters will be ordered along a developmental dimension using an external criterion. Solutions for three statistical problems will be given: (1) an algorithm that can handle large data sets and only renders non-degenerate clusters; (2) a goodness of fit test that is not affected by the fact that the number of possible response vectors by far out-weights the number of observed response vectors; and (3) a new technique, data expunction, that can be used to evaluate goodness-of-fit tests if the missing data mechanism is known.