Efficient software clustering technique using an adaptive and preventive dendrogram cutting approach

Efficient software clustering technique using an adaptive and preventive dendrogram cutting approach
复制标题

DOI:
10.1016/j.infsof.2013.07.002
复制
发表时间:
2013-11
期刊:
Inf. Softw. Technol.
影响因子:
--
通讯作者:
Chun Yong Chong;S. Lee;T. Ling
Chun Yong Chong;S. Lee;T. Ling
中科院分区:
其他
文献类型:
--
作者:
Chun Yong Chong;S. Lee;T. Ling

文献摘要

被引文献

相似文献

上下文软件集群是逆向工程中的一项关键技术,用于在资源有限的情况下恢复软件的高层抽象。非常有限的研究已经明确讨论了在设计中找到最佳的聚类集的问题,以及如何惩罚在clustering.<$This试图通过引入一个补充机制,以提高现有的凝聚聚类算法的单例集群的形成。为了解决架构恢复问题,所提出的方法侧重于减少冗余的努力和惩罚的形成,在聚类过程中的单例集群,同时保持的integrity of the results.MethodAn自动化的解决方案切割的树状图,是基于最小二乘回归,以找到最佳的切割水平。树状图是一个树形图,显示了软件实体集群的分类关系。此外,本文还引入了一个惩罚因子来惩罚形成单态的簇。在两个开源项目上进行了模拟。所提出的方法进行了比较,对详尽的和最高的差距树状图切割方法,以及两个著名的聚类有效性指标,即邓恩的指数和戴维斯-Bouldin index.ResultsWhen比较我们的聚类结果对原始的包图,我们的方法实现了平均准确率为90.07%,从两个模拟后,效用类被删除。源代码中的实用程序类由于其无所不在的行为而影响软件聚类的准确性。所提出的方法也成功地惩罚单例集群的形成clustering. ConclusionThe评价表明,所提出的方法可以提高质量的聚类结果,通过指导软件维护人员通过切割点的选择过程。该方法可以作为一种补充机制,以提高现有的聚类算法的有效性。
ContextSoftware clustering is a key technique that is used in reverse engineering to recover a high-level abstraction of the software in the case of limited resources. Very limited research has explicitly discussed the problem of finding the optimum set of clusters in the design and how to penalize for the formation of singleton clusters during clustering.ObjectiveThis paper attempts to enhance the existing agglomerative clustering algorithms by introducing a complementary mechanism. To solve the architecture recovery problem, the proposed approach focuses on minimizing redundant effort and penalizing for the formation of singleton clusters during clustering while maintaining the integrity of the results.MethodAn automated solution for cutting a dendrogram that is based on least-squares regression is presented in order to find the best cut level. A dendrogram is a tree diagram that shows the taxonomic relationships of clusters of software entities. Moreover, a factor to penalize clusters that will form singletons is introduced in this paper. Simulations were performed on two open-source projects. The proposed approach was compared against the exhaustive and highest gap dendrogram cutting methods, as well as two well-known cluster validity indices, namely, Dunn’s index and the Davies-Bouldin index.ResultsWhen comparing our clustering results against the original package diagram, our approach achieved an average accuracy rate of 90.07% from two simulations after the utility classes were removed. The utility classes in the source code affect the accuracy of the software clustering, owing to its omnipresent behavior. The proposed approach also successfully penalized the formation of singleton clusters during clustering.ConclusionThe evaluation indicates that the proposed approach can enhance the quality of the clustering results by guiding software maintainers through the cutting point selection process. The proposed approach can be used as a complementary mechanism to improve the effectiveness of existing clustering algorithms.