Sensitivity as a Complexity Measure for Sequence Classification Tasks

Sensitivity as a Complexity Measure for Sequence Classification Tasks
复制标题

DOI:
10.1162/tacl_a_00403
复制
发表时间:
2021-04
影响因子:
10.9
通讯作者:
Michael Hahn;Dan Jurafsky;Richard Futrell
Michael Hahn;Dan Jurafsky;Richard Futrell
中科院分区:
人文科学1区
文献类型:
--
作者:
Michael Hahn;Dan Jurafsky;Richard Futrell

文献摘要

被引文献

相似文献

摘要:我们使用布尔函数敏感性理论的新颖扩展,介绍了一个用于理解和预测序列分类任务复杂性的理论框架。给定输入序列的分布,函数的灵敏度量化了输入序列的不相交子集的数量,每个子集都可以单独更改以改变输出。我们认为标准序列分类方法偏向于学习低灵敏度函数,因此需要高灵敏度的任务更加困难。为此,我们通过分析证明简单的词汇分类器只能表达有限敏感度的函数,并且通过经验证明低敏感度函数对于 LSTM 来说更容易学习。然后,我们估计了 15 个 NLP 任务的敏感性,发现在 GLUE 中收集的具有挑战性的任务的敏感性高于简单文本分类任务,并且该敏感性预测了简单词汇分类器和没有预训练上下文嵌入的普通 BiLSTM 的性能。在任务中,灵敏度可以预测哪些输入对于此类简单模型来说是困难的。我们的结果表明,大规模预训练的上下文表示的成功部分源于它们提供了低灵敏度解码器可以从中提取信息的表示。
Abstract We introduce a theoretical framework for understanding and predicting the complexity of sequence classification tasks, using a novel extension of the theory of Boolean function sensitivity. The sensitivity of a function, given a distribution over input sequences, quantifies the number of disjoint subsets of the input sequence that can each be individually changed to change the output. We argue that standard sequence classification methods are biased towards learning low-sensitivity functions, so that tasks requiring high sensitivity are more difficult. To that end, we show analytically that simple lexical classifiers can only express functions of bounded sensitivity, and we show empirically that low-sensitivity functions are easier to learn for LSTMs. We then estimate sensitivity on 15 NLP tasks, finding that sensitivity is higher on challenging tasks collected in GLUE than on simple text classification tasks, and that sensitivity predicts the performance both of simple lexical classifiers and of vanilla BiLSTMs without pretrained contextualized embeddings. Within a task, sensitivity predicts which inputs are hard for such simple models. Our results suggest that the success of massively pretrained contextual representations stems in part because they provide representations from which information can be extracted by low-sensitivity decoders.