SEWordSim: software-specific word similarity database

SEWordSim: software-specific word similarity database
复制标题

DOI:
10.1145/2591062.2591071
复制
发表时间:
2014-05
期刊:
Companion Proceedings of the 36th International Conference on Software Engineering
影响因子:
--
通讯作者:
Yuan Tian;D. Lo;J. Lawall
Yuan Tian;D. Lo;J. Lawall
中科院分区:
其他
文献类型:
--
作者:
Yuan Tian;D. Lo;J. Lawall

文献摘要

被引文献

相似文献

测量单词的相似度对于准确表示和比较文档非常重要,从而改善许多自然语言处理 (NLP) 任务的结果。 NLP 社区提出了基于 WordNet 的各种测量方法,WordNet 是一个词汇数据库,包含许多单词对之间的关​​系。最近,人们提出了许多技术来解决软件工程问题,例如需要理解自然语言文档的代码搜索和故障定位,并且单词相似度的度量可以改善其结果。然而,WordNet 仅包含有关通用对话中的词义的信息,这通常与软件工程环境中的词义不同,并且已开发的特定于软件的单词相似性资源依赖于仅包含有限范围的单词和单词用法的数据源。在最近的工作中,我们提出了一种基于从 StackOverflow 自动收集的信息的单词相似度资源。我们发现该资源的结果按 3 点李克特量表给出的分数比基于 WordNet 的资源的结果高出 50% 以上。在这篇演示论文中,我们回顾了我们的数据收集方法,并提出了一个 Java API,以使生成的单词相似性资源在实践中有用。 SEWordSim 数据库和相关信息可以在 http://goo.gl/BVEAs8 找到。演示视频可在 http://goo.gl/dyNwyb 上获取。
Measuring the similarity of words is important in accurately representing and comparing documents, and thus improves the results of many natural language processing (NLP) tasks. The NLP community has proposed various measurements based on WordNet, a lexical database that contains relationships between many pairs of words. Recently, a number of techniques have been proposed to address software engineering issues such as code search and fault localization that require understanding natural language documents, and a measure of word similarity could improve their results. However, WordNet only contains information about words senses in general-purpose conversation, which often differ from word senses in a software-engineering context, and the software-specific word similarity resources that have been developed rely on data sources containing only a limited range of words and word uses. In recent work, we have proposed a word similarity resource based on information collected automatically from StackOverflow. We have found that the results of this resource are given scores on a 3-point Likert scale that are over 50% higher than the results of a resource based on WordNet. In this demo paper, we review our data collection methodology and propose a Java API to make the resulting word similarity resource useful in practice. The SEWordSim database and related information can be found at http://goo.gl/BVEAs8. Demo video is available at http://goo.gl/dyNwyb.