Building siamese attention-augmented recurrent convolutional neural networks for document similarity scoring

Building siamese attention-augmented recurrent convolutional neural networks for document similarity scoring
复制标题

DOI:
10.1016/j.ins.2022.10.032
复制
发表时间:
2022-10
期刊:
Inf. Sci.
影响因子:
--
通讯作者:
Sifei Han;Lingyun Shi;Russell Richie;F. Tsui
Sifei Han;Lingyun Shi;Russell Richie;F. Tsui
中科院分区:
其他
文献类型:
--
作者:
Sifei Han;Lingyun Shi;Russell Richie;F. Tsui

文献摘要

被引文献

相似文献

自动测量文档相似性在自然语言处理中是必不可少的,其应用范围从推荐到重复文档检测。文档相似性的最新方法通常涉及深度神经网络,但很少有关于如何组合不同架构的研究。因此,我们引入了结合了多种神经网络架构的连体注意力增强递归卷积神经网络(S-ARCNN)。在S-ARCNN的每个子网络中,文档通过双向长短期记忆(bi-LSTM)层,该层将表示发送到本地和全局文档模块。本地文档模块使用卷积、池化和注意层,而全局文档模块使用bi-LSTM的最后状态。局部和全局特征被连接起来形成一个单一的文档表示。使用Quora Question Pairs数据集,我们评估了S-ARCNN、Siamese卷积神经网络(S-CNN)、Siamese LSTM和两个BERT模型。虽然S-CNN(82.02%F1)总体上优于S-ARCNN(79.83%F1),但S-ARCNN在超过50个单词的重复问题对上的表现略优于S-CNN(39.96% vs. 39.42%准确率)。利用S-ARCNN在处理较长文档方面的潜在优势,S-ARCNN可以帮助研究人员识别具有相似研究兴趣的合作者,帮助编辑找到潜在的审稿人,或者将简历与职位描述相匹配。
Automatically measuring document similarity is imperative in natural language processing, with applications ranging from recommendation to duplicate document detection. State-of-the-art approach in document similarity commonly involves deep neural networks, yet there is little study on how different architectures may be combined. Thus, we introduce the Siamese Attention-augmented Recurrent Convolutional Neural Network (S-ARCNN) that combines multiple neural network architectures. In each subnetwork of S-ARCNN, a document passes through a bidirectional Long Short-Term Memory (bi-LSTM) layer, which sends representations to local and global document modules. A local document module uses convolution, pooling, and attention layers, whereas a global document module uses last states of the bi-LSTM. Both local and global features are concatenated to form a single document representation. Using the Quora Question Pairs dataset, we evaluated S-ARCNN, Siamese convolutional neural networks (S-CNNs), Siamese LSTM, and two BERT models. While S-CNNs (82.02% F1) outperformed S-ARCNN (79.83% F1) overall, S-ARCNN slightly outperformed S-CNN on duplicate question pairs with more than 50 words (39.96% vs. 39.42% accuracy). With the potential advantage of S-ARCNN for processing longer documents, S-ARCNN may help researchers identify collaborators with similar research interests, help editors find potential reviewers, or match resumes with job descriptions.