A hybrid approach of Weighted Fine-Tuned BERT extraction with deep Siamese Bi - LSTM model for semantic text similarity identification.

A hybrid approach of Weighted Fine-Tuned BERT extraction with deep Siamese Bi - LSTM model for semantic text similarity identification.
复制标题

加权微调的BERT提取的混合方法与深层Siamese BI -LSTM模型,用于语义文本相似性识别。

DOI:
10.1007/s11042-021-11771-6
复制
发表时间:
2022
影响因子:
3.6
通讯作者:
Revathy S
Revathy S
中科院分区:
计算机科学4区
文献类型:
--
作者:
Viji D;Revathy S

文献摘要

参考文献

被引文献

相似文献

传统的语义文本相似度方法需要大量的训练标注数据和人工干预。通常,它忽略了上下文信息和语序信息,从而导致了数据稀疏性问题和纬度爆炸问题。最近,深度学习方法被用来确定文本相似度。因此,本研究调查了自然语言处理应用任务在问题对或文档的文本相似度检测中的使用情况,并探索了相似度得分预测。提出了一种新的加权精调BERT特征提取与暹罗BiLSTM模型相结合的方法。该技术用于使用Quora数据集中的语义文本相似度来确定问题对集。利用BERT过程提取文本特征,然后嵌入加权的词。特征和权值被表示为嵌入向量,受到暹罗网络的不同层的影响。输入文本特征的嵌入向量采用深度暹罗BiLSTM模型在不同的层次进行训练。最后,为每个句子确定相似度得分,并学习语义文本相似度。通过与现有文本相似度检测方法的比较,从准确率、查准率、F1评分数据和召回值参数等方面对该框架进行了性能评估。与其他已有算法相比,该框架在确定语义文本相似度方面具有更高的效率和91%的准确率。
The conventional semantic text-similarity methods requires high amount of trained labeled data and also human interventions. Generally, it neglects the contextual-information and word-orders information resulted in data sparseness problem and latitudinal-explosion issue. Recently, deep-learning methods are used for determining text-similarity. Hence, this study investigates NLP application tasks usage in detecting text-similarity of question pairs or documents and explores the similarity score predictions. A new hybridized approach using Weighted Fine-Tuned BERT Feature extraction with Siamese Bi-LSTM model is implemented. The technique is employed for determining question pair sets using Semantic-text-similarity from Quora dataset. The text features are extracted using BERT process, followed by words embedding with weights. The features along with weight values, are represented as embedded vectors, are subjected to various layers of Siamese Networks. The embedded vectors of input text features were trained by using Deep Siamese Bi-LSTM model, in various layers. Finally, similarity scores are determined for each sentence, and the semantic text-similarity is learned. The performance evaluation of proposed-framework is established with respect to accuracy rate, precision value, F1 score data and Recall values parameters compared with other existing text-similarity detection methods. The proposed-framework exhibited higher efficiency rate with 91% in accuracy level in determining semantic-text-similarity compared with other existing algorithms.
DOI: 10.1016/j.suscom.2018.06.002
发表时间: 2018-09-01
影响因子: 4.5
作者:
Kumar, Mohit;Sharma, S. C.
通讯作者: Sharma, S. C.
DOI: 10.3390/app10175841
发表时间: 2020-09-01
影响因子: 2.7
作者:
Jang, Beakcheol;Kim, Myeonghwi;Kim, Jong Wook
通讯作者: Kim, Jong Wook
DOI: 10.1016/j.ipm.2017.01.002
发表时间: 2017-05-01
影响因子: 8.6
作者:
Al-Smadi, Mohammad;Jaradat, Zain;Jararweh, Yaser
通讯作者: Jararweh, Yaser
DOI: 10.1016/j.knosys.2019.07.013
发表时间: 2019-10-15
影响因子: 8.8
作者:
Nguyen, Hien T.;Duong, Phuc H.;Cambria, Erik
通讯作者: Cambria, Erik
将局部和全局特征组合到连体网络中以实现句子相似性
DOI: 10.1109/access.2020.2988918
发表时间: 2020-01-01
期刊: IEEE ACCESS
影响因子: 3.9
作者:
Li, Yulong;Zhou, Dong;Zhao, Wenyu
通讯作者: Zhao, Wenyu