Multi-objective Topic Modeling for Exploratory Search in Tech News

Multi-objective Topic Modeling for Exploratory Search in Tech News
复制标题

用于科技新闻探索性搜索的多目标主题建模

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
K. Vorontsov
K. Vorontsov
中科院分区:
--
文献类型:
--
作者:
A. Ianina;Lev Golitsyn;K. Vorontsov

文献摘要

被引文献

相似文献

探索性搜索是信息检索的一种范式,用户的目的是更好地了解主题领域。要做到这一点,用户需要多次重复与搜索引擎的“查询-浏览-优化”交互。我们考虑由长文本查询制定的典型探索性搜索任务。人们通常在半小时内解决这样的任务,并使用传统的搜索工具迭代地找到数十个文档。本文的目标是在不影响搜索质量的前提下,将耗时的多步骤过程简化为一步。概率主题建模是检索与长文本查询在语义上相关的文档的一种合适的文本挖掘技术。我们使用主题模型的加性正则化(ARTM)来构建一个满足多个目标的模型。模型应该具有稀疏、多样和可解释的主题。此外,它还应该包含元数据和多模态数据,如n-grams、作者、标签和类别。平衡正则化准则是ARTM的一个重要问题。我们利用坐标优化技术自动选择正则化轨迹来解决这一问题。我们使用了开源库BigARTM的并行在线实现。我们的评估技术是基于众包的,包括评估人员的两个任务:手动探索性搜索和明确的相关性反馈。在两种流行的科技新闻媒体上的实验表明,我们基于主题的探索性搜索优于评估器和简单的基线,达到了85-92%的准确率和召回率。
Exploratory search is a paradigm of information retrieval, in which the user’s intention is to learn the subject domain better. To do this the user repeats “query–browse–refine” interactions with the search engine many times. We consider typical exploratory search tasks formulated by long text queries. People usually solve such a task in about half an hour and find dozens of documents using conventional search facilities iteratively. The goal of this paper is to reduce the time-consuming multi-step process to one step without impairing the quality of the search. Probabilistic topic modeling is a suitable text mining technique to retrieve documents, which are semantically relevant to a long text query. We use the additive regularization of topic models (ARTM) to build a model that meets multiple objectives. The model should have sparse, diverse and interpretable topics. Also, it should incorporate meta-data and multimodal data such as n-grams, authors, tags and categories. Balancing the regularization criteria is an important issue for ARTM. We tackle this problem with coordinate-wise optimization technique, which chooses the regularization trajectory automatically. We use the parallel online implementation of ARTM from the open source library BigARTM. Our evaluation technique is based on crowdsourcing and includes two tasks for assessors: the manual exploratory search and the explicit relevance feedback. Experiments on two popular tech news media show that our topic-based exploratory search outperforms assessors as well as simple baselines, achieving precision and recall of about 85–92%.