An Attention-based Deep Relevance Model for Few-shot Document Filtering
An Attention-based Deep Relevance Model for Few-shot Document Filtering
复制标题
一种基于注意力的深度相关性模型,用于小样本文档过滤
DOI:
10.1145/3419972
复制
发表时间:
2020-10
期刊:
影响因子:
--
通讯作者:
Chen Haiqing
中科院分区:
文献类型:
--
作者:
Liu Bulou;Li Chenliang;Zhou Wei;Ji Feng;Duan Yu;Chen Haiqing
With the large quantity of textual information produced on the Internet, a critical necessity is to filter out the irrelevant information and organize the rest into categories of interest (e.g., an emerging event). However, supervised-learning document filtering methods heavily rely on a large number of labeled documents for model training. Manually identifying plenty of positive examples for each category is expensive and time-consuming. Also, it is unrealistic to cover all the categories from an evolving text source that covers diverse kinds of events, user opinions, and daily life activities. In this article, we propose a novel attention-based deep relevance model for few-shot document filtering (named ADRM), inspired by the relevance feedback methodology proposed for ad hoc retrieval. ADRM calculates the relevance score between a document and a category by taking a set of seed words and a few seed documents relevant to the category. It constructs the category-specific conceptual representation of the document based on the corresponding seed words and seed documents. Specifically, to filter irrelevant yet noisy information in the seed documents, ADRM employs two types of attention mechanisms (namely whole-match attention and max-match attention) and generates category-specific representations for them. Then ADRM is devised to extract the relevance signals by modeling the hidden feature interactions in the word embedding space. The relevance signals are extracted through a gated convolutional process, a self-attention layer, and a relevance aggregation layer. Extensive experiments on three real-world datasets show that ADRM consistently outperforms the existing technical alternatives, including the conventional classification and retrieval baselines, and the state-of-the-art deep relevance ranking models for few-shot document filtering. We also perform an ablation study to demonstrate that each component in ADRM is effective for enhancing filtering performance. Further analysis shows that ADRM is robust under varying parameter settings.
登录
查看更多内容
DOI:
--
发表时间:
2014-12
期刊:
ArXiv
影响因子:
--
作者:
Baotian Hu;Zhengdong Lu;Hang Li;Qingcai Chen
通讯作者:
Baotian Hu;Zhengdong Lu;Hang Li;Qingcai Chen
DOI:
--
发表时间:
2012-11
期刊:
--
影响因子:
--
作者:
John R. Frank;Max Kleiman-Weiner;D. Roberts;Feng Niu;Ce Zhang;Christopher Ré;I. Soboroff
通讯作者:
John R. Frank;Max Kleiman-Weiner;D. Roberts;Feng Niu;Ce Zhang;Christopher Ré;I. Soboroff
DOI:
10.1145/2983323.2983728
发表时间:
2016-09
期刊:
Proceedings of the 25th ACM International on Conference on Information and Knowledge Management
影响因子:
--
作者:
R. Reinanda;E. Meij;M. de Rijke
通讯作者:
R. Reinanda;E. Meij;M. de Rijke
DOI:
--
发表时间:
2014-09
期刊:
--
影响因子:
--
作者:
Yaroslav Ganin;V. Lempitsky
通讯作者:
Yaroslav Ganin;V. Lempitsky
DOI:
10.1609/aaai.v29i1.9506
发表时间:
2015-01
期刊:
--
影响因子:
--
作者:
Xingyuan Chen;Yunqing Xia;Peng Jin;John A. Carroll
通讯作者:
Xingyuan Chen;Yunqing Xia;Peng Jin;John A. Carroll