Comparing Text Representations: A Theory-Driven Approach

Comparing Text Representations: A Theory-Driven Approach
复制标题

DOI:
10.18653/v1/2021.emnlp-main.449
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Gregory Yauney;David M. Mimno
Gregory Yauney;David M. Mimno
中科院分区:
其他
文献类型:
--
作者:
Gregory Yauney;David M. Mimno

文献摘要

相似文献

当代NLP的大部分进展都来自于学习表示,例如掩蔽语言模型(MLM)上下文嵌入,它将具有挑战性的问题转化为简单的分类任务。但我们如何量化和解释这种影响呢?我们采用计算学习理论的通用工具来适应文本数据集的特定特征,并提出了一种评估表示和任务之间兼容性的方法。尽管许多任务可以很容易地用简单的词袋(BOW)表示来解决,但BOW在困难的自然语言推理任务上表现不佳。对于一个这样的任务,我们发现BOW不能区分真实的和随机的标签,而预先训练的MLM表示在真实的和随机标签之间的区别比BOW大72倍。该方法提供了基于分类的NLP任务的难度的校准的定量测量,使得能够在表示之间进行比较,而不需要可能对初始化和超参数敏感的经验评估。该方法提供了一个新的角度来看,在数据集中的模式和这些模式与特定的标签对齐。
Much of the progress in contemporary NLP has come from learning representations, such as masked language model (MLM) contextual embeddings, that turn challenging problems into simple classification tasks. But how do we quantify and explain this effect? We adapt general tools from computational learning theory to fit the specific characteristics of text datasets and present a method to evaluate the compatibility between representations and tasks. Even though many tasks can be easily solved with simple bag-of-words (BOW) representations, BOW does poorly on hard natural language inference tasks. For one such task we find that BOW cannot distinguish between real and randomized labelings, while pre-trained MLM representations show 72x greater distinction between real and random labelings than BOW. This method provides a calibrated, quantitative measure of the difficulty of a classification-based NLP task, enabling comparisons between representations without requiring empirical evaluations that may be sensitive to initializations and hyperparameters. The method provides a fresh perspective on the patterns in a dataset and the alignment of those patterns with specific labels.