A novel selective learning based transformer encoder architecture with enhanced word representation

A novel selective learning based transformer encoder architecture with enhanced word representation
复制标题

DOI:
10.1007/s10489-022-03865-x
复制
发表时间:
2022-08
影响因子:
5.3
通讯作者:
Wazib Ansar;Saptarsi Goswami;A. Chakrabarti;B. Chakraborty
Wazib Ansar;Saptarsi Goswami;A. Chakrabarti;B. Chakraborty
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wazib Ansar;Saptarsi Goswami;A. Chakrabarti;B. Chakraborty

文献摘要

相似文献

随着具有注意机制的转换器的出现,自然语言处理(NLP)的进步是多方面的。然而,这些模型具有巨大的复杂性和巨大的计算开销。此外,这些模型的性能依赖于对输入文本进行编码的特征表示策略。针对这些问题,我们提出了一种新的变换编码器结构,该结构具有选择性学习遗忘网络(SLFN)和通过词性特征嵌入增强上下文单词表示(PSCE)。新的SLFN通过门控机制选择性地保留文本中的重要信息。它支持并行处理,捕获远程依赖关系,同时在处理长序列时提高转换器的效率。而直观的PSCE处理多义词,根据语境区分词的屈折变化,并有效地理解语篇中的句法和语义信息。与BERT相比,单块架构的参数减少了96.1%,效率极高。与传统的变压器结构相比,该结构的准确率提高了6.8%,并且在对来自不同领域的三个数据集进行情感分析时,与各种最先进的模型相比有了显著的改进。
With the advent of transformers having attention mechanisms, the advancements in Natural Language Processing (NLP) have been manifold. However, these models possess huge complexity and enormous computational overhead. Besides, the performance of such models relies on the feature representation strategy for encoding the input text. To address these issues, we propose a novel transformer encoder architecture with Selective Learn-Forget Network (SLFN) and contextualized word representation enhanced through Parts-of-Speech Characteristics Embedding (PSCE). The novel SLFN selectively retains significant information in the text through a gated mechanism. It enables parallel processing, captures long-range dependencies and simultaneously increases the transformer’s efficiency while processing long sequences. While the intuitive PSCE deals with polysemy, distinguishes word-inflections based on context and effectively understands the syntactic as well as semantic information in the text. The single-block architecture is extremely efficient with 96.1% reduced parameters compared to BERT. The proposed architecture yields 6.8% higher accuracy than vanilla transformer architecture and appreciable improvement over various state-of-the-art models for sentiment analysis over three data-sets from diverse domains.