The need for software specific natural language techniques

The need for software specific natural language techniques
复制标题

对软件特定自然语言技术的需求

DOI:
10.1007/s10664-017-9566-5
复制
发表时间:
2018
影响因子:
4.1
通讯作者:
Christopher Morrell
Christopher Morrell
中科院分区:
计算机科学2区
文献类型:
--
作者:
D. Binkley;Dawn J Lawrie;Christopher Morrell

文献摘要

被引文献

相似文献

二十多年来,软件工程(SE)研究人员一直在从信息检索(IR)中引入工具和技术。初步结果相当积极。例如,当应用于诸如功能定位或重新建立可追溯性链接之类的问题时,IR技术本身就能很好地工作,并且通常与更传统的源代码分析技术(如静态和动态分析)相结合甚至更好。然而,最近有越来越多的意识,在SE研究人员,IR工具和技术的设计工作在不同的假设下比那些持有的软件系统。因此,考虑专门设计用于软件的IR启发的工具和技术可能是有益的。这项工作的目的之一是提供定量的经验证据来支持这一观察。为了做到这一点,引入了一种新的技术,这种技术可以捕捉到信息需求中的困难程度,也就是信息需求者希望知道的真实的、往往是潜在的信息。新技术用于比较两个领域:自然语言(NL)和SE。对数据的分析得出了三个重要的发现。首先,SE信息需求的难度分布的变化不同于NL信息需求;其次,收集年龄在NL集合之间的差异中起作用;最后,所使用的检索模型对结果的影响不大。
For over two decades, software engineering (SE) researchers have been importing tools and techniques from information retrieval (IR). Initial results have been quite positive. For example, when applied to problems such as feature location or re-establishing traceability links, IR techniques work well on their own, and often even better in combination with more traditional source code analysis techniques such as static and dynamic analysis. However, recently there has been growing awareness among SE researchers that IR tools and techniques are designed to work under different assumptions than those that hold for a software system. Thus it may be beneficial to consider IR-inspired tools and techniques that are specifically designed to work with software. One aim of this work is to provide quantitative empirical evidence in support of this observation. To do so a new technique is introduced that captures the level of difficulty found in an information need, the true, often latent, information that a searcher desires to know. The new technique is used to compare two domains: Natural Language (NL) and SE. Analysis of the data leads to three significant findings. First, the variation in the distribution of difficulty of the SE information needs differs from that of the NL information needs; second, collection age plays a role in the differences between the NL collections; and finally, the retrieval model used has little impact on the results.