Diversity-Based Generalization for Neural Unsupervised Text Classification under Domain Shift

Diversity-Based Generalization for Neural Unsupervised Text Classification under Domain Shift
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Jitin Krishnan;Hemant Purohit;H. Rangwala
Jitin Krishnan;Hemant Purohit;H. Rangwala
中科院分区:
其他
文献类型:
--
作者:
Jitin Krishnan;Hemant Purohit;H. Rangwala

文献摘要

被引文献

相似文献

领域适应方法寻求从源域学习,并将其推广到看不见的目标领域。目前,针对主观文本分类问题的领域自适应方法都是半监督的,并且使用未标记的目标数据和已标记的源数据。本文基于一种简单而有效的基于多样性的泛化思想,不需要未标记的目标数据,提出了一种新的单任务文本分类问题的领域自适应方法。多样性通过迫使模型不依赖相同的特征进行预测,从而促进模型更好地泛化和不分青红皂白地进行领域转换。我们将这一概念应用于神经网络中最容易解释的组件--关注层。为了产生足够的多样性,我们创建了一个多头注意模型,并在注意头之间注入了多样性约束,使得每个头学习的方式不同。我们进一步扩展了我们的模型,通过三次训练并设计了一个过程,在三次训练的分类器的注意力头部之间增加了多样性约束。使用亚马逊评论的标准基准数据集和新构建的危机事件数据集进行的广泛评估表明,我们的完全无监督方法与竞争的半监督基线相匹配。我们的结果表明,确保足够多样性的机器学习体系结构可以更好地推广;鼓励未来的研究在不使用未标记的目标数据的情况下设计普遍可用的学习模型。
Domain adaptation approaches seek to learn from a source domain and generalize it to an unseen target domain. At present, the state-of-the-art domain adaptation approaches for subjective text classification problems are semi-supervised; and use unlabeled target data along with labeled source data. In this paper, we propose a novel method for domain adaptation of single-task text classification problems based on a simple but effective idea of diversity-based generalization that does not require unlabeled target data. Diversity plays the role of promoting the model to better generalize and be indiscriminate towards domain shift by forcing the model not to rely on same features for prediction. We apply this concept on the most explainable component of neural networks, the attention layer. To generate sufficient diversity, we create a multi-head attention model and infuse a diversity constraint between the attention heads such that each head will learn differently. We further expand upon our model by tri-training and designing a procedure with an additional diversity constraint between the attention heads of the tri-trained classifiers. Extensive evaluation using the standard benchmark dataset of Amazon reviews and a newly constructed dataset of Crisis events shows that our fully unsupervised method matches with the competing semi-supervised baselines. Our results demonstrate that machine learning architectures that ensure sufficient diversity can generalize better; encouraging future research to design ubiquitously usable learning models without using unlabeled target data.