Sameness: An Experiment in Code Search

Sameness: An Experiment in Code Search
复制标题

相同性:代码搜索的实验

DOI:
--
复制
发表时间:
2015
期刊:
2015 IEEE/ACM 12th Working Conference on Mining Software Repositories
影响因子:
--
通讯作者:
A. Hoek
A. Hoek
中科院分区:
--
文献类型:
--
作者:
Lee Martie;A. Hoek

文献摘要

被引文献

相似文献

到目前为止,大多数专用代码搜索引擎使用排名算法,只关注查询和结果之间的相关性。在实践中,这意味着开发人员可能会收到全部来自同一项目的搜索结果,全部使用相同的外部库实现相同的算法,或者全部表现出相同的复杂性或大小,以及其他不太理想的可能性。在本文中,我们提出,代码搜索引擎还应该找到不同的和简洁(简短但完整)的代码结果集。我们提出了四种新的算法,使用相关性,多样性和简洁性排名代码搜索结果。为了评估这些算法以及多样性和简洁性在代码搜索中的价值,21名专业程序员被要求比较竞争算法产生的前十名结果。我们发现,我们的两个新算法产生的十大结果是强烈的程序员的首选。
To date, most dedicated code search engines use ranking algorithms that focus only on the relevancy between the query and the results. In practice, this means that a developer may receive search results that are all drawn from the same project, all implement the same algorithm using the same external library, or all exhibit the same complexity or size, among other possibilities that are less than ideal. In this paper, we propose that code search engines should also locate both diverse and concise (brief but complete) sets of code results. We present four novel algorithms that use relevance, diversity, and conciseness in ranking code search results. To evaluate these algorithms and the value of diversity and conciseness in code search, twenty-one professional programmers were asked to compare pairs of top ten results produced by competing algorithms. We found that two of our new algorithms produce top ten results that are strongly preferred by the programmers.