QScored: A Large Dataset of Code Smells and Quality Metrics

QScored: A Large Dataset of Code Smells and Quality Metrics
复制标题

DOI:
10.1109/msr52588.2021.00080
复制
发表时间:
2021-05
期刊:
2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR)
影响因子:
--
通讯作者:
Tushar Sharma;Marouane Kessentini
Tushar Sharma;Marouane Kessentini
中科院分区:
其他
文献类型:
--
作者:
Tushar Sharma;Marouane Kessentini

文献摘要

被引文献

相似文献

代码质量方面,如代码气味和代码质量度量被广泛用于探索性和经验性的软件工程研究。在这些研究中,研究人员花费大量的时间和精力,不仅选择适当的主题系统,而且还分析它们以收集所需的代码质量信息。在本文中,我们介绍了QScored数据集;该数据集包含超过8.6万个C#和Java GitHub存储库的代码质量信息,其中包含超过11亿行代码。代码质量信息包含7种检测到的架构气味,20种设计气味,11种实现气味,以及在项目,包,类和方法级别计算的27个常用代码质量度量。数据集的可用性将通过使大量活跃的GitHub存储库随时可用信息来促进涉及代码质量方面的实证研究。
Code quality aspects such as code smells and code quality metrics are widely used in exploratory and empirical software engineering research. In such studies, researchers spend a substantial amount of time and effort to not only select the appropriate subject systems but also to analyze them to collect the required code quality information. In this paper, we present QScored dataset; the dataset contains code quality information of more than 86 thousand C# and Java GitHub repositories containing more than 1.1 billion lines of code. The code quality information contains seven kinds of detected architecture smells, 20 kinds of design smells, eleven kinds of implementation smells, and 27 commonly used code quality metrics computed at project, package, class, and method levels. Availability of the dataset will facilitate empirical studies involving code quality aspects by making the information readily available for a large number of active GitHub repositories.