Undersampling Improves Hypernymy Prototypicality Learning

Undersampling Improves Hypernymy Prototypicality Learning
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Koki Washio;Tsuneaki Kato
Koki Washio;Tsuneaki Kato
中科院分区:
其他
文献类型:
--
作者:
Koki Washio;Tsuneaki Kato

文献摘要

相似文献

本文主要研究基于未登录词对分布表示的有监督的上位词检测方法。Levy等人。(2015)证明了有监督的上位词检测存在训练数据中过多的上位词(fiting Hypernym)。我们发现,在这项任务上出现fi过多的问题是由数据集的一个特性引起的,这种特性源于所使用的语言资源的内在结构,即层次叙词表。本文提出的简单的数据预处理方法缓解了这一问题。更准确地说,我们通过实验证明了训练数据中Hypernymy Classifi超过fit上位词的问题源于词表的准树结构带来的词频分布不对称,并提出了一种简单的基于词频的欠采样方法,该方法可以有效地缓解Overfi设置,提高未知词对的分布原型学习。
This paper focuses on supervised hypernymy detection using distributional representations for unknown word pairs. Levy et al. (2015) demonstrated that supervised hypernymy detection suffers from overfitting hypernyms in training data. We show that the problem of overfitting on this task is caused by a characteristic of datasets, which stems from the inherent structure of the language resources used, hierarchical thesauri. The simple data preprocessing method proposed in this paper alleviates this problem. To be more precise, we demonstrate through experiments that the problem that hypernymy classifiers overfit hypernyms in training data comes from a skewed word frequency distribution brought by the quasi-tree structure of a thesaurus, which is a major resource of lexical semantic relation data, and propose a simple undersampling method based on word frequencies that can effectively alleviate overfitting and improve distributional prototypicality learning for unknown word pairs.