Experiences on Clustering High-Dimensional Data using pbdR

Experiences on Clustering High-Dimensional Data using pbdR
复制标题

使用pbdR对高维数据进行聚类的经验

DOI:
10.1145/3144763.3144768
复制
发表时间:
2017
期刊:
Proceedings of the 1st International Workshop on Software Engineering for High Performance Computing in Computational and Data-enabled Science & Engineering
影响因子:
--
通讯作者:
Mockus, Audris
Mockus, Audris
中科院分区:
--
文献类型:
--
作者:
Amreen, Sadika;Mockus, Audris

文献摘要

参考文献

相似文献

针对高性能计算(HPC)环境的软件工程[1],尤其是针对大数据的软件工程[5]面临着一系列独特的挑战,包括中间件和计算环境的高度复杂性。因此,使科学家更容易利用高性能计算的工具至关重要。我们提供了一个使用这种高效中间件pbdR[9]的经验报告,它允许科学家使用R编程语言,而至少在名义上不必掌握许多层的HPC基础设施,如OpenMPI[4]和ScalaPACK[2]。目的评估中间件在多大程度上帮助提高科学家的生产力,我们使用pbdR来解决我们作为科学家正在研究的一个实际问题。我们的大数据来自GitHub和其他项目托管网站上的提交,我们正在尝试根据这些提交消息的文本对开发人员进行集群。上下文我们需要能够识别每次提交的开发人员,并识别单个开发人员的提交。提交中的开发人员标识符如登录、电子邮件和姓名通常有多种拼写方式,因为这些信息可能来自不同的版本控制系统(Git、Mercurial、SVN等)并且可能取决于使用哪台计算机(在主文件夹的.git/config中指定的)。方法我们训练Doc2Vec[7]模型,其中将现有凭证用作文档标识符,然后使用所得到的230万个标识符的200维向量来对这些标识符进行集群,从而每个集群代表一个特定的个体。距离矩阵占用32TB,因此,总体上是高性能计算,特别是PBDR的良好目标。PbdR允许数据分布在计算节点上,甚至在PMClust包中实现了K-均值和混合模型聚类技术。结果我们使用战略原型[3]来评估pbdR的能力,发现a)中间件的使用需要广泛了解其内部工作原理,从而否定了许多预期的好处;b)所实现的算法不适合n、p和k(样本大小、数据维度和聚类数量)的特定组合;C)基于批处理作业的开发环境大大增加了开发时间。结论除了Basili等人的发现外,我们还发现高性能计算基础设施及其开发环境的实施质量对开发生产率有很大影响。
MotivationSoftware engineering for High Performace Computing (HPC) environments in general [1] and for big data in particular [5] faces a set of unique challenges including high complexity of middleware and of computing environments. Tools that make it easier for scientists to utilize HPC are, therefore, of paramount importance. We provide an experience report of using one of such highly effective middleware pbdR [9] that allow the scientist to use R programming language without, at least nominally, having to master many layers of HPC infrastructure, such as OpenMPI [4] and ScalaPACK [2].Objectiveto evaluate the extent to which middleware helps improve scientist productivity, we use pbdR to solve a real problem that we, as scientists, are investigating. Our big data comes from the commits on GitHub and other project hosting sites and we are trying to cluster developers based on the text of these commit messages.ContextWe need to be able to identify developer for every commit and to identify commits for a single developer. Developer identifiers in the commits, such as login, email, and name are often spelled in multiple ways since that information may come from different version control systems (Git, Mercurial, SVN, ...) and may depend on which computer is used (what is specified in .git/config of the home folder).MethodWe train Doc2Vec [7] model where existing credentials are used as a document identifier and then use the resulting 200-dimensional vectors for the 2.3M identifiers to cluster these identifiers so that each cluster represents a specific individual. The distance matrix occupies 32TB and, therefore, is a good target for HPC in general and pbdR in particular. pbdR allows data to be distributed over computing nodes and even has implemented K-means and mixture-model clustering techniques in the package pmclust.ResultsWe used strategic prototyping [3] to evaluate the capabilities of pbdR and discovered that a) the use of middleware required extensive understanding of its inner workings thus negating many of the expected benefits; b) the implemented algorithms were not suitable for the particular combination of n, p, and k (sample size, data dimension, and the number of clusters); c) the development environment based on batch jobs increases development time substantially.ConclusionsIn addition to finding from Basili et al., we find that the quality of the implementation of HPC infrastructure and its development environment has a tremendous effect on development productivity.
大数据系统的软件工程
DOI: --
发表时间: 2016
期刊: IEEE Software
影响因子: 3.3
作者:
I. Gorton;A. Bener;A. Mockus
通讯作者: A. Mockus
Le Fort I截骨固定方法及术后骨片移位检查
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
藤尾正人;佐世暁;荻須宏太;土屋周平;酒井陽;日比英晴
通讯作者: 日比英晴