Doc2Vec-based Approach for Extracting Diverse Evaluation Expressions from Online Review Data

Doc2Vec-based Approach for Extracting Diverse Evaluation Expressions from Online Review Data
复制标题

DOI:
10.1145/3487664.3487773
复制
发表时间:
2021-11
期刊:
The 23rd International Conference on Information Integration and Web Intelligence
影响因子:
--
通讯作者:
Kosuke Kurihara;Yoshiyuki Shoji;Sumio Fujita;M. Dürst
Kosuke Kurihara;Yoshiyuki Shoji;Sumio Fujita;M. Dürst
中科院分区:
其他
文献类型:
--
作者:
Kosuke Kurihara;Yoshiyuki Shoji;Sumio Fujita;M. Dürst

文献摘要

相似文献

提出了一种从在线影评文本中提取特定关键字查询的多样表达的方法。当人们看一部让他们哭的电影时,他们通常不会说“我哭了”。相反,他们会用一些委婉的语言,比如“我需要一条手帕”或者“我的妆花了”。为了使用评论文本实现基于观众反应的信息检索,例如“让我哭的电影”,必须收集各种释义表达以用于任意查询。我们提出的方法通过对Doc 2 Vec应用两个扩展来从评论数据集中提取这样的表达式:1)它改变训练句子的粒度以减轻上下文的缺乏,以及2)它提前应用查询扩展进行相似性计算。我们进行了一个大规模的实验,使用众包与129万实际句子从雅虎!日本的电影。实验结果表明,改变训练数据粒度和增加查询扩展都是有效的,以准确地收集更多样化的表达式,具有类似的给定查询的含义。
This paper proposes a method for extracting diverse expressions from online movie review texts for a given keyword query. When people watch a movie that makes them cry, they generally do not say “I cried.” Instead, they use such euphemistic language as “I needed a handkerchief” or “My makeup was running.” To enable information retrieval based on audience reactions such as “movies that make me cry” using review texts, a variety of paraphrased expressions must be collected for arbitrary queries. Our proposed method extracts such expressions from review datasets by applying two extensions to Doc2Vec: 1) it changes the granularity of the training sentences to mitigate a lack of context, and 2) it applies query expansion for similarity calculation in advance. We conducted a large-scale experiment using crowdsourcing with 1.29 million actual sentences taken from Yahoo! Movies, Japan. The experimental result revealed that changing the training data granularity and adding the query expansion are both effective to accurately collect more diverse expressions that have a meaning similar to the given query.