A dataset for evaluating identifier splitters

A dataset for evaluating identifier splitters
复制标题

用于评估标识符分割器的数据集

DOI:
10.1109/msr.2013.6624055
复制
发表时间:
2013
期刊:
2013 10th Working Conference on Mining Software Repositories (MSR)
影响因子:
--
通讯作者:
K. Vijay
K. Vijay
中科院分区:
--
文献类型:
--
作者:
D. Binkley;Dawn J Lawrie;L. Pollock;Emily Hill;K. Vijay

文献摘要

被引文献

相似文献

软件工程和演变技术最近已开始利用源代码中的自然语言信息。这样做的关键步骤是将标识符分配到其组成词中。概念简单,但标识符分裂引发了几个具有挑战性的问题,从而导致了一系列分裂技术。因此,研究界将受益于促进标识符分裂技术比较研究的数据集(即黄金集)。从8,522个个人分裂判断中构建了一组2,663个拆分标识符,可以从www.cs.loyola.edu/~binkley/ludiso获得。描述了该集合的构建和旨在有效使用的观察结果。
Software engineering and evolution techniques have recently started to exploit the natural language information in source code. A key step in doing so is splitting identifiers into their constituent words. While simple in concept, identifier splitting raises several challenging issues, leading to a range of splitting techniques. Consequently, the research community would benefit from a dataset (i.e., a gold set) that facilitates comparative studies of identifier splitting techniques. A gold set of 2,663 split identifiers was constructed from 8,522 individual human splitting judgements and can be obtained from www.cs.loyola.edu/~binkley/ludiso. This set's construction and observations aimed at its effective use are described.