Automatic clustering constraints derivation from object-oriented software using weighted complex network with graph theory analysis

Automatic clustering constraints derivation from object-oriented software using weighted complex network with graph theory analysis
复制标题

DOI:
10.1016/j.jss.2017.08.017
复制
发表时间:
2017-11
期刊:
J. Syst. Softw.
影响因子:
--
通讯作者:
Chun Yong Chong;S. Lee
Chun Yong Chong;S. Lee
中科院分区:
其他
文献类型:
--
作者:
Chun Yong Chong;S. Lee

文献摘要

被引文献

相似文献

约束聚类或半监督聚类由于其灵活性,可以结合领域专家或辅助信息的最小监督来帮助改善经典无监督聚类技术的聚类结果,因此受到了很多关注。在软件重新模块化领域,经典的无监督软件集群技术已被证明有助于恢复文档或设计不完善的软件系统的软件设计的高级抽象。然而,缺乏出于相同目的集成约束集群来帮助提高软件系统的模块化性的工作。然而,由于时间和预算的限制,对于具有软件先验知识的领域专家来说,审查每个软件工件并按需提供监督是费力且不现实的。我们的目标是通过提出一种自动化方法来填补这一研究空白,该方法基于对所分析软件的图论分析,从软件系统的隐式结构中导出聚类约束。对40个开源面向对象软件系统进行的评估表明,所提出的方法可以作为在不存在领域专家的情况下导出聚类约束的替代解决方案,从而有助于提高聚类结果的整体准确性。
Constrained clustering or semi-supervised clustering has received a lot of attention due to its flexibility of incorporating minimal supervision of domain experts or side information to help improve clustering results of classic unsupervised clustering techniques. In the domain of software remodularisation, classic unsupervised software clustering techniques have proven to be useful to aid in recovering a high-level abstraction of the software design of poorly documented or designed software systems. However, there is a lack of work that integrates constrained clustering for the same purpose to help improve the modularity of software systems. Nevertheless, due to time and budget constraints, it is laborious and unrealistic for domain experts who have prior knowledge about the software to review each and every software artifact and provide supervision on an on-demand basis. We aim to fill this research gap by proposing an automated approach to derive clustering constraints from the implicit structure of software system based on graph theory analysis of the analysed software. Evaluations conducted on 40 open-source object-oriented software systems show that the proposed approach can serve as an alternative solution to derive clustering constraints in situations where domain experts are non-existent, thus helping to improve the overall accuracy of clustering results.