Code and Named Entity Recognition in StackOverflow

Code and Named Entity Recognition in StackOverflow
复制标题

DOI:
10.18653/v1/2020.acl-main.443
复制
发表时间:
2020-05
期刊:
--
影响因子:
--
通讯作者:
Jeniya Tabassum;Mounica Maddela;Wei Xu;Alan Ritter
Jeniya Tabassum;Mounica Maddela;Wei Xu;Alan Ritter
中科院分区:
其他
文献类型:
--
作者:
Jeniya Tabassum;Mounica Maddela;Wei Xu;Alan Ritter

文献摘要

相似文献

随着大量的编程文本在互联网上变得容易获得,人们对一起学习自然语言和计算机代码的兴趣越来越大。例如,StackOverflow目前有超过1500万个编程相关的问题,由850万用户编写。与此同时,仍然缺乏基本的NLP技术来识别出现在自然语言句子中的代码令牌或与软件相关的命名实体。在本文中,我们介绍了一个新的命名实体识别(NER)语料库的计算机编程领域,由15,372个句子与20个细粒度的实体类型注释。我们在来自StackOverflow的1.52亿个句子上训练了域内BERT表示(BERTOverflow),这使得F1分数比现成的BERT绝对增加了+10。我们还提出了SoftNER模型,该模型在StackOverflow数据上的代码和命名实体识别方面获得了79.10 F-1的总体得分。我们的SoftNER模型采用了一个上下文无关的代码令牌分类器与语料库级别的功能,以改善基于BERT的标记模型。我们的代码和数据可在https://github.com/jeniyat/StackOverflowNER/上获得
There is an increasing interest in studying natural language and computer code together, as large corpora of programming texts become readily available on the Internet. For example, StackOverflow currently has over 15 million programming related questions written by 8.5 million users. Meanwhile, there is still a lack of fundamental NLP techniques for identifying code tokens or software-related named entities that appear within natural language sentences. In this paper, we introduce a new named entity recognition (NER) corpus for the computer programming domain, consisting of 15,372 sentences annotated with 20 fine-grained entity types. We trained in-domain BERT representations (BERTOverflow) on 152 million sentences from StackOverflow, which lead to an absolute increase of +10 F1 score over off-the-shelf BERT. We also present the SoftNER model which achieves an overall 79.10 F-1 score for code and named entity recognition on StackOverflow data. Our SoftNER model incorporates a context-independent code token classifier with corpus-level features to improve the BERT-based tagging model. Our code and data are available at: https://github.com/jeniyat/StackOverflowNER/