A novel selective learning based transformer encoder architecture with enhanced word representation
A novel selective learning based transformer encoder architecture with enhanced word representation
复制标题
DOI:
10.1007/s10489-022-03865-x
复制
发表时间:
2022-08
影响因子:
5.3
通讯作者:
Wazib Ansar;Saptarsi Goswami;A. Chakrabarti;B. Chakraborty
中科院分区:
文献类型:
--
作者:
Wazib Ansar;Saptarsi Goswami;A. Chakrabarti;B. Chakraborty
With the advent of transformers having attention mechanisms, the advancements in Natural Language Processing (NLP) have been manifold. However, these models possess huge complexity and enormous computational overhead. Besides, the performance of such models relies on the feature representation strategy for encoding the input text. To address these issues, we propose a novel transformer encoder architecture with Selective Learn-Forget Network (SLFN) and contextualized word representation enhanced through Parts-of-Speech Characteristics Embedding (PSCE). The novel SLFN selectively retains significant information in the text through a gated mechanism. It enables parallel processing, captures long-range dependencies and simultaneously increases the transformer’s efficiency while processing long sequences. While the intuitive PSCE deals with polysemy, distinguishes word-inflections based on context and effectively understands the syntactic as well as semantic information in the text. The single-block architecture is extremely efficient with 96.1% reduced parameters compared to BERT. The proposed architecture yields 6.8% higher accuracy than vanilla transformer architecture and appreciable improvement over various state-of-the-art models for sentiment analysis over three data-sets from diverse domains.