MUDABlue: an automatic categorization system for open source repositories

MUDABlue: an automatic categorization system for open source repositories
复制标题

DOI:
10.1016/j.jss.2005.06.044
复制
发表时间:
2004-11
期刊:
11th Asia-Pacific Software Engineering Conference
影响因子:
--
通讯作者:
Shinji Kawaguchi;P. Garg;M. Matsushita;Katsuro Inoue
Shinji Kawaguchi;P. Garg;M. Matsushita;Katsuro Inoue
中科院分区:
其他
文献类型:
--
作者:
Shinji Kawaguchi;P. Garg;M. Matsushita;Katsuro Inoue

文献摘要

被引文献

相似文献

开放源码社区通常使用软件存储库来归档各种软件项目及其源代码、邮件列表讨论、文档、错误报告等。例如,SourceForge目前托管着超过7万个开源软件系统。由于丰富的信息内容的大小,这样的存储库提供了许多在项目之间共享信息的机会。例如,人们想知道一组彼此相关或相似的项目,以便项目组可以协作和共享他们的工作。然而,由于典型存储库中有数千个项目,手动查找相关项目可能很困难。因此,我们提出了MUDABlue,一个自动对软件系统进行分类的工具。MUDABlue有三个主要方面:1)它不依赖于源代码以外的其他信息,2)它自动确定类别集,3)它允许一个软件系统成为多个类别的成员。MUDABlue有一个Web界面来可视化确定的类别,这使得浏览软件存储库变得容易。我们通过将MUDABlue生成的类别与其他一些现有研究工具的分类能力进行比较,证明了MUDABlue的分类能力的有效性。
Open source communities typically use a software repository to archive various software projects with their source code, mailing list discussions, documentation, bug reports, and so forth. For example, SourceForge currently hosts over seventy thousand open source software systems. Because of the size of the rich information content, such repositories offer numerous opportunities for sharing information among projects. For example, one would like to know a set of projects that are related or similar to each other, so that the project groups can collaborate and share their work. With thousands of projects in typical repositories, however, manually locating related projects can be difficult. Hence, we propose MUDABlue, a tool that automatically categorizes software systems. MUDABlue has three major aspects: 1) it relies on no other information than the source code, 2) it determines category sets automatically, and 3) it allows a software system to be a member of multiple categories. MUDABlue has a Web interface to visualize determined categories, which eases browsing a software repository. We show the effectiveness of MUDABlue's categorization capability by comparing its generated categories with that of some other existing research tools.