Vec2Node: Self-Training with Tensor Augmentation for Text Classification with Few Labels

Vec2Node: Self-Training with Tensor Augmentation for Text Classification with Few Labels
复制标题

DOI:
10.1007/978-3-031-26390-3_33
复制
发表时间:
2022
影响因子:
14.9
通讯作者:
S. Abdali;Subhabrata Mukherjee;E. Papalexakis
S. Abdali;Subhabrata Mukherjee;E. Papalexakis
中科院分区:
生物学2区
文献类型:
--
作者:
S. Abdali;Subhabrata Mukherjee;E. Papalexakis

文献摘要

相似文献

在深度神经网络等最先进的机器学习模型中,最近的进展严重依赖于大量的标记训练数据,而这些数据对于许多应用来说是很难获得的。为了解决标签稀缺的问题,最近的工作集中在数据增强技术上,以创建合成训练数据。在这项工作中,我们提出了一种新的数据增强方法,利用张量分解来生成合成样本,利用文本中的局部和全局信息,减少概念漂移。我们开发了Vec2Node,它利用来自域内未标记数据的自我训练,并使用张紧词嵌入来增强,这比最先进的模型有了显著的改进,特别是在低资源环境下。例如,仅使用标记的训练数据,Vec2Node通过以下方式提高基本模型的精度。此外,Vec2Node利用张量嵌入生成可解释的扩展数据。
Recent advances in state-of-the-art machine learning models like deep neural networks heavily rely on large amounts of labeled training data which is difficult to obtain for many applications. To address label scarcity, recent work has focused on data augmentation techniques to create synthetic training data. In this work, we propose a novel approach of data augmentation leveraging tensor decomposition to generate synthetic samples by exploiting local and global information in text and reducing concept drift. We developVec2Nodethat leverages self-training from in-domain unlabeled data augmented with tensorized word embeddings that significantly improves over state-of-the-art models, particularly in low-resource settings. For instance, with onlyof labeled training data,Vec2Nodeimproves the accuracy of a base model by. Furthermore,Vec2Nodegenerates explicable augmented data leveraging tensor embeddings.