Rissanen Data Analysis: Examining Dataset Characteristics via Description Length

Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
复制标题

DOI:
--
复制
发表时间:
2021-03
期刊:
--
影响因子:
--
通讯作者:
Ethan Perez;Douwe Kiela;Kyunghyun Cho
Ethan Perez;Douwe Kiela;Kyunghyun Cho
中科院分区:
其他
文献类型:
--
作者:
Ethan Perez;Douwe Kiela;Kyunghyun Cho

文献摘要

被引文献

相似文献

我们介绍了一种方法来确定某些功能是否有助于实现给定数据的准确模型。我们将标签视为由由具有不同功能的子例程组成的程序从输入中生成的,并且我们认为,当且仅当调用它的最小程序比没有的最小程序时,子例程才有用。由于最小程序长度是不可兼容的,因此我们将标签的最小描述长度(MDL)视为代理,从而为我们提供了一种分析数据集特性的理论基础方法。我们将MDL之父称为Rissanen数据分析(RDA)方法,我们在NLP的各种环境中展示了其适用性,从评估在回答问题之前生成子问题的实用性到分析理性和理性的价值解释,调查语音不同部分的重要性,并发现数据集性别偏见。
We introduce a method to determine if a certain capability helps to achieve an accurate model of given data. We view labels as being generated from the inputs by a program composed of subroutines with different capabilities, and we posit that a subroutine is useful if and only if the minimal program that invokes it is shorter than the one that does not. Since minimum program length is uncomputable, we instead estimate the labels' minimum description length (MDL) as a proxy, giving us a theoretically-grounded method for analyzing dataset characteristics. We call the method Rissanen Data Analysis (RDA) after the father of MDL, and we showcase its applicability on a wide variety of settings in NLP, ranging from evaluating the utility of generating subquestions before answering a question, to analyzing the value of rationales and explanations, to investigating the importance of different parts of speech, and uncovering dataset gender bias.