CodeDJ: Reproducible Queries over Large-Scale Software Repositories (Artifact)

CodeDJ: Reproducible Queries over Large-Scale Software Repositories (Artifact)
复制标题

CodeDJ:大规模软件存储库的可重复查询(工件)

DOI:
--
复制
发表时间:
2021
期刊:
Dagstuhl Artifacts Ser.
影响因子:
--
通讯作者:
J. Vitek
J. Vitek
中科院分区:
--
文献类型:
--
作者:
Petr Maj;Konrad Siek;A. Kovalenko;J. Vitek

文献摘要

被引文献

相似文献

分析大量的代码库是现代软件工程研究的主要内容--这是诸如GitHub这样的大型软件库出现的一个受欢迎的副作用。选择一个人应该分析的项目是一个劳动密集型的过程,如果选择不能代表感兴趣的人群,这个过程可能会导致有偏见的结果。研究人员面临的一个问题是,软件存储库公开的界面只允许最基本的查询。Code DJ是用于查询存储库的基础设施,该存储库由持久性数据存储(使用从GitHub获取的数据不断更新)和带有Rust查询接口的内存数据库组成。代码DJ支持重现性,使用数据存储的过去状态确定地回答历史查询;因此,研究人员可以重现发布的结果。为了说明代码DJ的好处,我们确定了一项已发表研究的数据中的偏差,并通过使用新数据重复分析,证明了该研究的结论对项目的选择是敏感的。
Analyzing massive code bases is a staple of modern software engineering research – a welcome side-effect of the advent of large-scale software repositories such as GitHub. Selecting which projects one should analyze is a labor-intensive process, and a process that can lead to biased results if the selection is not representative of the population of interest. One issue faced by researchers is that the interface exposed by software repositories only allows the most basic of queries. Code DJ is an infrastructure for querying repositories composed of a persistent datastore, constantly updated with data acquired from GitHub, and an in-memory database with a Rust query interface. Code DJ supports reproducibility, historical queries are answered deterministically using past states of the datastore; thus researchers can reproduce published results. To illustrate the benefits of Code DJ , we identify biases in the data of a published study and, by repeating the analysis with new data, we demonstrate that the study’s conclusions were sensitive to the choice of projects.