Implementation of Self-organizing Maps with Python

Implementation of Self-organizing Maps with Python
复制标题

用Python实现自组织映射

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Li Yuan
Li Yuan
中科院分区:
--
文献类型:
--
作者:
Li Yuan

文献摘要

被引文献

相似文献

自组织映射(Self-Organizing Maps,SOM)作为人工神经网络的一员,自20世纪80年代以来得到了广泛的研究,并已在C、Fortran、R [1]和Python [2]中实现。Python是一种高效的高级语言,在机器学习领域被广泛使用多年,但大多数用Python编写的SOM相关软件包只执行模型构建和可视化。然而,用R编写的POPSOM包能够执行模型构建和可视化之外的功能,例如使用统计方法评估模型的质量和绘制神经元的边缘概率分布。为了给Python用户提供POPSOM包的优势,将POPSOM包迁移为基于Python的包是很重要的。这项研究显示了这种实施的细节。具体实现有三个主要任务:1)将POPSOM包从R语言迁移到Python语言; 2)将源代码从过程化编程范式重构为面向对象编程范式; 3)通过在模型构造函数中添加规范化选项来改进包。除了用Python构建模型外,还嵌入了Fortran,以显着加快模型构建的速度。最后的程序已经完成,需要保证程序的正确性。实现这一目标的最佳方法是将基于Python的程序的输出与基于R的程序生成的输出进行比较。对于模型构造函数,SOM算法在初始阶段随机选取神经元的权向量,然后在训练过程中随机选取输入向量。由于这两个随机因素,人们不能期望相同的输入(数据集)会产生完全相同的输出(神经元)。相反,为了证明Python程序正常工作,在这个项目中提出并应用了两种解决方案:1)测量分别由R和Python函数生成的两个神经元之间的向量的平均差; 2)测量两个神经元的方差比和特征均值差。除了模型构建之外,模型可视化和其他以神经元作为输入的功能应该通过馈送相同的输入(神经元)返回相同的结果。上述验证的详细内容将在以下章节中介绍。
As a member of Artificial Neural Networks, Self-Organizing Maps (SOMs) have been well researched since 1980s, and have been implemented in C, Fortran, R [1] and Python [2]. Python is an efficient high-level language widely used in the machine learning field for years, but most of the SOM-related packages which are written in Python only perform model construction and visualization. However, the POPSOM package, written in R, is capable of performing functionality beyond model construction and visualization, such as evaluating the model’s quality with statistical methods and plotting marginal probability distributions of the neurons. In order to give the Python user the POPSOM package’s advantages, it is important to migrate the POPSOM package to be Python-based. This study shows the details of this implementation. There are three major tasks for the implementation: 1) Migrate the POPSOM package from R to Python; 2) Refactor the source code from procedural programming paradigm to object-oriented programming paradigm; 3) Improve the package by adding normalization options to the model construction function. In addition to constructing the model in Python, Fortran is also embedded to accelerate the speed of model construction significantly in this project. The final program has been completed, and it is necessary to guarantee the correctness of the program. The best way to achieve this goal is to compare the output of the Python-based program to the output generated by the R-based program. For the model construction function, the SOM algorithm initializes the weight vector of the neurons randomly at the very beginning, and then selects the input vectors randomly during the training. Due to these two random factors, one cannot expect the same input (data set) will result in exactly the same output (neurons). Instead, to give evidence that the Python program is working properly, there are two solutions that have been proposed and applied in this project: 1) measuring the average difference of vectors between two neurons which have been generated by the R and Python functions respectively; 2) measuring the ratio of the variances and the difference of features’ mean for the two neurons. Besides the model construction, model visualization and other functions which take neurons as their input should return the same results by feeding the same input (neurons). The detail of above verification will be represented in the following chapters.