Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability

Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability
复制标题

DOI:
10.18653/v1/p17-1075
复制
发表时间:
2017-07
期刊:
--
影响因子:
--
通讯作者:
Saku Sugawara;Yusuke Kido;Hikaru Yokono;Akiko Aizawa
Saku Sugawara;Yusuke Kido;Hikaru Yokono;Akiko Aizawa
中科院分区:
其他
文献类型:
--
作者:
Saku Sugawara;Yusuke Kido;Hikaru Yokono;Akiko Aizawa

文献摘要

相似文献

了解阅读理解(RC)数据集的质量对于自然语言理解系统的开发非常重要。在这项研究中,两个类的指标被用来评估RC数据集:先决条件的技能和可读性。我们将这些类应用于六个现有的数据集,包括MCTest和SQuAD,并根据每个指标和两个类之间的相关性突出了数据集的特征。我们的数据集分析表明,RC数据集的可读性不会直接影响问题的难度,并且可以创建一个易于阅读但难以回答的RC数据集。
Knowing the quality of reading comprehension (RC) datasets is important for the development of natural-language understanding systems. In this study, two classes of metrics were adopted for evaluating RC datasets: prerequisite skills and readability. We applied these classes to six existing datasets, including MCTest and SQuAD, and highlighted the characteristics of the datasets according to each metric and the correlation between the two classes. Our dataset analysis suggests that the readability of RC datasets does not directly affect the question difficulty and that it is possible to create an RC dataset that is easy to read but difficult to answer.