Training and Prediction Data Discrepancies: Challenges of Text Classification with Noisy, Historical Data

Training and Prediction Data Discrepancies: Challenges of Text Classification with Noisy, Historical Data
复制标题

训练和预测数据差异:使用嘈杂的历史数据进行文本分类的挑战

DOI:
--
复制
发表时间:
2018
期刊:
NUT@EMNLP
影响因子:
--
通讯作者:
R. A. Kreek
R. A. Kreek
中科院分区:
--
文献类型:
--
作者:
Emilia Apostolova;R. A. Kreek

文献摘要

被引文献

相似文献

用于文本分类的行业数据集很少用于此目的。在大多数情况下,数据和目标预测是累积的历史数据的副产品,通常充满了噪声,存在于基于文本的文档以及目标标签中。在这项工作中,我们解决了在嘈杂的历史数据上计算的性能指标如何反映未来机器学习模型输入的性能的问题。结果证明了脏训练数据集的实用性,这些数据集用于为更清洁(和不同)的预测输入构建预测模型。
Industry datasets used for text classification are rarely created for that purpose. In most cases, the data and target predictions are a by-product of accumulated historical data, typically fraught with noise, present in both the text-based document, as well as in the targeted labels. In this work, we address the question of how well performance metrics computed on noisy, historical data reflect the performance on the intended future machine learning model input. The results demonstrate the utility of dirty training datasets used to build prediction models for cleaner (and different) prediction inputs.