High-dimensional cluster analysis with the masked EM algorithm.

High-dimensional cluster analysis with the masked EM algorithm.
复制标题

DOI:
10.1162/neco_a_00661
复制
发表时间:
2014-11
期刊:
影响因子:
2.9
通讯作者:
Harris KD
Harris KD
中科院分区:
计算机科学4区
文献类型:
--
作者:
Kadir SN;Goodman DF;Harris KD

文献摘要

被引文献

相似文献

聚类分析在高维中面临两个问题:第一,“维数灾难”,这可能导致过拟合和泛化性能差;第二,传统算法处理大量高维数据所需的时间太长。我们描述了这些问题的解决方案,设计用于下一代高通道数神经探针的“尖峰分选”的应用。在这个问题中,只有一个小的特征子集提供关于任何一个数据向量的聚类成员身份的信息,但是这个信息特征子集对于所有数据点是不相同的,使得经典的特征选择无效。我们引入了一个“屏蔽EM”算法,允许在数千个维度上对多达数百万个点进行准确和时间效率高的聚类。我们证明了它的适用性,合成数据,并在现实世界中的高通道计数尖峰排序数据。
Cluster analysis faces two problems in high dimensions: first, the “curse of dimensionality” that can lead to overfitting and poor generalization performance; and second, the sheer time taken for conventional algorithms to process large amounts of high-dimensional data. We describe a solution to these problems, designed for the application of “spike sorting” for next-generation high channel-count neural probes. In this problem, only a small subset of features provide information about the cluster member-ship of any one data vector, but this informative feature subset is not the same for all data points, rendering classical feature selection ineffective. We introduce a “Masked EM” algorithm that allows accurate and time-efficient clustering of up to millions of points in thousands of dimensions. We demonstrate its applicability to synthetic data, and to real-world high-channel-count spike sorting data.