Experiences on Clustering High-Dimensional Data using pbdR
Experiences on Clustering High-Dimensional Data using pbdR
复制标题
使用pbdR对高维数据进行聚类的经验
DOI:
10.1145/3144763.3144768
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Mockus, Audris
中科院分区:
文献类型:
--
作者:
Amreen, Sadika;Mockus, Audris
MotivationSoftware engineering for High Performace Computing (HPC) environments in general [1] and for big data in particular [5] faces a set of unique challenges including high complexity of middleware and of computing environments. Tools that make it easier for scientists to utilize HPC are, therefore, of paramount importance. We provide an experience report of using one of such highly effective middleware pbdR [9] that allow the scientist to use R programming language without, at least nominally, having to master many layers of HPC infrastructure, such as OpenMPI [4] and ScalaPACK [2].Objectiveto evaluate the extent to which middleware helps improve scientist productivity, we use pbdR to solve a real problem that we, as scientists, are investigating. Our big data comes from the commits on GitHub and other project hosting sites and we are trying to cluster developers based on the text of these commit messages.ContextWe need to be able to identify developer for every commit and to identify commits for a single developer. Developer identifiers in the commits, such as login, email, and name are often spelled in multiple ways since that information may come from different version control systems (Git, Mercurial, SVN, ...) and may depend on which computer is used (what is specified in .git/config of the home folder).MethodWe train Doc2Vec [7] model where existing credentials are used as a document identifier and then use the resulting 200-dimensional vectors for the 2.3M identifiers to cluster these identifiers so that each cluster represents a specific individual. The distance matrix occupies 32TB and, therefore, is a good target for HPC in general and pbdR in particular. pbdR allows data to be distributed over computing nodes and even has implemented K-means and mixture-model clustering techniques in the package pmclust.ResultsWe used strategic prototyping [3] to evaluate the capabilities of pbdR and discovered that a) the use of middleware required extensive understanding of its inner workings thus negating many of the expected benefits; b) the implemented algorithms were not suitable for the particular combination of n, p, and k (sample size, data dimension, and the number of clusters); c) the development environment based on batch jobs increases development time substantially.ConclusionsIn addition to finding from Basili et al., we find that the quality of the implementation of HPC infrastructure and its development environment has a tremendous effect on development productivity.
影响因子:
3.3
作者:
I. Gorton;A. Bener;A. Mockus
通讯作者:
A. Mockus
DOI:
--
发表时间:
2019
期刊:
影响因子:
--
作者:
藤尾正人;佐世暁;荻須宏太;土屋周平;酒井陽;日比英晴
通讯作者:
日比英晴