Modeling Personal Biases in Language Use by Inducing Personalized Word Embeddings

Modeling Personal Biases in Language Use by Inducing Personalized Word Embeddings
复制标题

DOI:
10.18653/v1/n19-1215
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Daisuke Oba;Naoki Yoshinaga;Shoetsu Sato;Satoshi Akasaki;Masashi Toyoda
Daisuke Oba;Naoki Yoshinaga;Shoetsu Sato;Satoshi Akasaki;Masashi Toyoda
中科院分区:
其他
文献类型:
--
作者:
Daisuke Oba;Naoki Yoshinaga;Shoetsu Sato;Satoshi Akasaki;Masashi Toyoda

文献摘要

被引文献

相似文献

个体语言使用存在偏差;相同的词(例如,凉爽)用于表达不同的含义(例如,温度范围)或不同的词(例如,阴天、朦胧)用于描述相同的含义。在这项研究中,我们提出了一种通过解决主观文本任务而获得的个性化词嵌入来对词义(以下称为语义变化)中的个人偏见进行建模的方法,同时将不同个人使用的词视为不同的词。为了防止个性化词嵌入受到其他不相关偏见的污染,我们解决了从给定评论中识别评论目标(客观输出)的任务。为了稳定这种极端多类分类的训练,我们通过元数据识别进行多任务学习。从 RateBeer 检索到的评论的实验结果证实,所获得的个性化词嵌入提高了情感分析以及目标任务的准确性。对获得的个性化词嵌入的分析揭示了与频繁词和形容词相关的语义变化趋势。
There exist biases in individual’s language use; the same word (e.g., cool) is used for expressing different meanings (e.g., temperature range) or different words (e.g., cloudy, hazy) are used for describing the same meaning. In this study, we propose a method of modeling such personal biases in word meanings (hereafter, semantic variations) with personalized word embeddings obtained by solving a task on subjective text while regarding words used by different individuals as different words. To prevent personalized word embeddings from being contaminated by other irrelevant biases, we solve a task of identifying a review-target (objective output) from a given review. To stabilize the training of this extreme multi-class classification, we perform a multi-task learning with metadata identification. Experimental results with reviews retrieved from RateBeer confirmed that the obtained personalized word embeddings improved the accuracy of sentiment analysis as well as the target task. Analysis of the obtained personalized word embeddings revealed trends in semantic variations related to frequent and adjective words.