SpecNFS: A Challenge Dataset Towards Extracting Formal Models from Natural Language Specifications

SpecNFS: A Challenge Dataset Towards Extracting Formal Models from Natural Language Specifications
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Sayontan Ghosh;Amanpreet Singh;Alex Merenstein;S. Smolka;E. Zadok;Niranjan Balasubramanian
Sayontan Ghosh;Amanpreet Singh;Alex Merenstein;S. Smolka;E. Zadok;Niranjan Balasubramanian
中科院分区:
其他
文献类型:
--
作者:
Sayontan Ghosh;Amanpreet Singh;Alex Merenstein;S. Smolka;E. Zadok;Niranjan Balasubramanian

文献摘要

相似文献

NLP可以帮助建立验证复杂系统的形式化模型吗?我们研究这个挑战的背景下,解析网络文件系统(NFS)规范。我们定义了一个语义依赖问题SpecIR,一种表示语言,我们引入模型的句子出现在NFS规范文档(RFC)作为IF-THEN语句,并提出了一个注释的数据集的1,198句。我们开发和评估这个问题的语义依赖分析系统。评估表明,即使使用最先进的语言模型,也有很大的改进空间,最好的模型在命名实体识别和依赖链接预测子任务中的F1得分分别仅为60.5和33.3。我们还发布其他未标记的数据和其他与域相关的文本。实验表明,这些额外的资源增加了F1的措施时,用于简单的领域适应和迁移学习为基础的方法,提出了富有成效的方向,为进一步的研究
Can NLP assist in building formal models for verifying complex systems? We study this challenge in the context of parsing Network File System (NFS) specifications. We define a semantic-dependency problem over SpecIR, a representation language we introduce to model sentences appearing in NFS specification documents (RFCs) as IF-THEN statements, and present an annotated dataset of 1,198 sentences. We develop and evaluate semantic-dependency parsing systems for this problem. Evaluations show that even when using a state-of-the-art language model, there is significant room for improvement, with the best models achieving an F1 score of only 60.5 and 33.3 in the named-entity-recognition and dependency-link-prediction sub-tasks, respectively. We also release additional unlabeled data and other domain-related texts. Experiments show that these additional resources increase the F1 measure when used for simple domain-adaption and transfer-learning-based approaches, suggesting fruitful directions for further research