Vec2Node: Self-Training with Tensor Augmentation for Text Classification with Few Labels
Vec2Node: Self-Training with Tensor Augmentation for Text Classification with Few Labels
复制标题
DOI:
10.1007/978-3-031-26390-3_33
复制
发表时间:
2022
影响因子:
14.9
通讯作者:
S. Abdali;Subhabrata Mukherjee;E. Papalexakis
中科院分区:
文献类型:
--
作者:
S. Abdali;Subhabrata Mukherjee;E. Papalexakis
Recent advances in state-of-the-art machine learning models like deep neural networks heavily rely on large amounts of labeled training data which is difficult to obtain for many applications. To address label scarcity, recent work has focused on data augmentation techniques to create synthetic training data. In this work, we propose a novel approach of data augmentation leveraging tensor decomposition to generate synthetic samples by exploiting local and global information in text and reducing concept drift. We developVec2Nodethat leverages self-training from in-domain unlabeled data augmented with tensorized word embeddings that significantly improves over state-of-the-art models, particularly in low-resource settings. For instance, with onlyof labeled training data,Vec2Nodeimproves the accuracy of a base model by. Furthermore,Vec2Nodegenerates explicable augmented data leveraging tensor embeddings.