Towards an automatic requirements classification in a new Spanish dataset

Towards an automatic requirements classification in a new Spanish dataset
复制标题

在新的西班牙语数据集中实现自动需求分类

DOI:
10.1109/re54965.2022.00039
复制
发表时间:
2022
期刊:
2022 IEEE 30th International Requirements Engineering Conference (RE)
影响因子:
--
通讯作者:
M. R. Luaces
M. R. Luaces
中科院分区:
--
文献类型:
--
作者:
María Isabel Limaylla Lunarejo;Nelly Condori;M. R. Luaces

文献摘要

被引文献

相似文献

机器学习(ML)算法已经成为软件需求分类的有力工具。然而,大多数研究集中在要求是在英语,与其他语言的关注较少。由于缺乏西班牙语的数据集,我们从拉科鲁尼亚大学最终学位项目的要求集合中创建了一个新的数据集。在本文中,我们研究了文本矢量化技术与ML算法的组合在西班牙语数据集中的需求分类中表现最好。我们发现,具有TF-IDF的SVM给出了最高的f1分数(功能和非功能分类分别为0.95和0.79)。
Machine Learning (ML) algorithms have become a powerful instrument in software requirements classification. Nevertheless, most of the research focusing on requirements is in English, with less attention to other languages. Given a lack of datasets in Spanish, we created a new dataset from a collection of requirements from final degree projects from the University of A Coruña. In this paper, we investigate which combinations of text vectorization techniques with ML algorithms perform best for requirements classification in a Spanish dataset. We found that SVM with TF-IDF gives the highest f1-score (0.95 and 0.79 for functional and non-functional classification).