Self-attention-based neural networks for refining the overlength product titles

Self-attention-based neural networks for refining the overlength product titles
复制标题

基于自注意力的神经网络用于细化超长的产品标题

DOI:
10.1007/s11042-021-10908-x
复制
发表时间:
2021-06
期刊:
Multim. Tools Appl.
影响因子:
--
通讯作者:
Yuming Lin
Yuming Lin
中科院分区:
其他
文献类型:
--
作者:
You Li;Guoyong Cai;Aoying Zhou;Yu Fu;Yuming Lin

文献摘要

参考文献

相似文献

在线卖家经常在电子商务平台上制作冗余和冗长的产品文本标题,并提供额外的信息以吸引客户的注意力。当在移动的应用上显示这些超长的产品标题时,它们成为问题。本文研究了如何对冗余和超长的产品标题进行提炼,以生成简洁、信息量大的标题。首先,通过预测原标题中的某个词是否会保留在最终短标题中,将长标题的精化问题转化为序列分类问题。然后,提出了一种基于自注意力的神经网络,从原始标题中提取信息量最大的词来构建简短的标题。所提出的基本模型还扩展了门控递归单元(GRU)神经网络和门控机制,以改善位置编码过程,并从不同方向学习编码特征的权重。在此基础上,设计了一种基于开放数据集LESD 4EC的冗余产品标题压缩分析数据集的构造算法。最后,在重建的数据集上进行了大量的实验,以证明所提出的方法的有效性和效率。实验结果表明,该方法在查准率、查全率、F1值和平均绝对误差、运行时间和空间开销等方面均优于现有方法.
Online sellers often produce redundant and lengthy product textual titles with extra information on e-commerce platforms to attract the attentions of customers. Such overlength product titles become a problem when they are displayed on mobile applications. In this paper, the problem of refining redundant and overlength product titles is studied to generate concise and informative titles. First, the task of refining the long title is transformed into a sequential classification problem by predicting whether a word in original title will remain in finial short title. Then, a self-attention-based neural network is proposed to extract the most informative words from original title to construct the short title. The proposed basic model is also extended with a gated recurrent unit (GRU) neural network and a gating mechanism to improve the position encoding process and learn the weights of encoding features from different directions. Moreover, an algorithm is designed to construct the datasets for redundant product title compression analysis based on the open dataset LESD4EC. Finally, extensive experiments are implemented on the rebuilt datasets to demonstrate the effectiveness and efficiency of the proposed methods. The experimental results show that the proposed methods significantly outperform the state-of-the-art methods based on the precision, recall,F1value and the mean absolute error, as well as runtime and space cost.
DOI: 10.1016/s0004-3702(02)00222-9
发表时间: 2002-07
期刊: Artif. Intell.
影响因子: --
作者:
Kevin Knight;D. Marcu
通讯作者: Kevin Knight;D. Marcu
DOI: 10.18653/v1/n19-2009
发表时间: 2019-04
期刊: --
影响因子: --
作者:
Jianguo Zhang;Pengcheng Zou;Zhao Li;Yao Wan;Xiuming Pan;Yu Gong;Philip S. Yu
通讯作者: Jianguo Zhang;Pengcheng Zou;Zhao Li;Yao Wan;Xiuming Pan;Yu Gong;Philip S. Yu
DOI: 10.1145/2483669.2483674
发表时间: 2013-06
期刊: ACM Trans. Intell. Syst. Technol.
影响因子: --
作者:
Trevor Cohn;Mirella Lapata
通讯作者: Trevor Cohn;Mirella Lapata
DOI: 10.1609/aaai.v33i01.33019460
发表时间: 2018-03
期刊: ArXiv
影响因子: --
作者:
Yu Gong;Xusheng Luo;Kenny Q. Zhu;Shichen Liu;Wenwu Ou
通讯作者: Yu Gong;Xusheng Luo;Kenny Q. Zhu;Shichen Liu;Wenwu Ou
DOI: 10.1145/3269206.3271722
发表时间: 2018-08
期刊: Proceedings of the 27th ACM International Conference on Information and Knowledge Management
影响因子: --
作者:
Fei Sun;Peng Jiang;Hanxiao Sun;Changhua Pei;Wenwu Ou;Xiaobo Wang
通讯作者: Fei Sun;Peng Jiang;Hanxiao Sun;Changhua Pei;Wenwu Ou;Xiaobo Wang