Quantifying the Evaluation of Heuristic Methods for Textual Data Augmentation

Quantifying the Evaluation of Heuristic Methods for Textual Data Augmentation
复制标题

DOI:
10.18653/v1/2020.wnut-1.26
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Omid Kashefi;R. Hwa
Omid Kashefi;R. Hwa
中科院分区:
其他
文献类型:
--
作者:
Omid Kashefi;R. Hwa

文献摘要

被引文献

相似文献

数据增强已被证明可以有效地为机器学习提供更多的训练数据,并产生更健壮的分类器。然而,对于某些问题,可能会有多个增强启发式,选择使用哪一个可能会显著影响训练的成功。在这项工作中,我们提出了一个衡量增强启发式的指标;具体来说,我们通过考虑不同类别的增强样本分布之间的差异来量化示例“难以区分”的程度。在两个预测任务(积极/消极情绪和冗长/简洁)中进行多重启发式实验,通过揭示不同类别的分布差异与分类精度之间的联系,验证了我们的说法。
Data augmentation has been shown to be effective in providing more training data for machine learning and resulting in more robust classifiers. However, for some problems, there may be multiple augmentation heuristics, and the choices of which one to use may significantly impact the success of the training. In this work, we propose a metric for evaluating augmentation heuristics; specifically, we quantify the extent to which an example is “hard to distinguish” by considering the difference between the distribution of the augmented samples of different classes. Experimenting with multiple heuristics in two prediction tasks (positive/negative sentiment and verbosity/conciseness) validates our claims by revealing the connection between the distribution difference of different classes and the classification accuracy.