Beyond a Bag of Words: Using PULSAR to Extract Judgments on Specific Human Rights at Scale

Beyond a Bag of Words: Using PULSAR to Extract Judgments on Specific Human Rights at Scale
复制标题

DOI:
10.1515/peps-2018-0030
复制
发表时间:
2018-10
期刊:
Peace Economics, Peace Science and Public Policy
影响因子:
--
通讯作者:
Baekkwan Park;Michael Colaresi;Kevin Greene
Baekkwan Park;Michael Colaresi;Kevin Greene
中科院分区:
其他
文献类型:
--
作者:
Baekkwan Park;Michael Colaresi;Kevin Greene

文献摘要

被引文献

相似文献

情绪、判断和表达的立场是国际关系和社会科学中的重要概念。然而,当代的定量研究传统上避免了最直接和最微妙的信息来源:政治和社会文本。相比之下,定性研究长期以来一直依赖于文本中的模式来了解公众舆论,社会问题,国际联盟的条款和政治家的立场的详细趋势。然而,定性的人类阅读并不能与当前可用的加速增长的大量数字信息相匹配。研究人员需要能够从文本中提取有意义的观点和判断的自动化工具。因此,有一个新兴的机会,结合基于模型的,推理的定量方法的重点,如理想点模型,高分辨率,定性的语言和立场的解释。我们建议,使用替代简单的词袋(BOW)表示和重新关注方面的情感表示的文本将有助于研究人员系统地提取人们的判断和正在判断的规模。下面的实验结果表明,我们的方法自动提取方面和情感MWE对,在分类任务中优于BOW,同时提供更多可解释的参数。通过连接表达的情感和被判断的方面,PULSAR(将非结构化语言解析为情感-方面表示)也对理解问题位置的潜在维度和用文本估计的理想点有着深刻的影响。我们将文本解析为方面的方法-情感表达恢复了表达性短语(类似于分类投票)以及正在判断的方面(类似于账单)。因此,PULSAR或类似的未来系统,为在现有的理想点模型中系统分析高维意见和判断开辟了新的途径。
Abstract Sentiment, judgments and expressed positions are crucial concepts across international relations and the social sciences more generally. Yet, contemporary quantitative research has conventionally avoided the most direct and nuanced source of this information: political and social texts. In contrast, qualitative research has long relied on the patterns in texts to understand detailed trends in public opinion, social issues, the terms of international alliances, and the positions of politicians. Yet, qualitative human reading does not scale to the accelerating mass of digital information available currently. Researchers are in need of automated tools that can extract meaningful opinions and judgments from texts. Thus, there is an emerging opportunity to marry the model-based, inferential focus of quantitative methodology, as exemplified by ideal point models, with high resolution, qualitative interpretations of language and positions. We suggest that using alternatives to simple bag of words (BOW) representations and re-focusing on aspect-sentiment representations of text will aid researchers in systematically extracting people’s judgments and what is being judged at scale. The experimental results below show that our approach which automates the extraction of aspect and sentiment MWE pairs, outperforms BOW in classification tasks, while providing more interpretable parameters. By connecting expressed sentiment and the aspects being judged, PULSAR (Parsing Unstructured Language into Sentiment-Aspect Representations) also has deep implications for understanding the underlying dimensionality of issue positions and ideal points estimated with text. Our approach to parsing text into aspects-sentiment expressions recovers both expressive phrases (akin to categorical votes), as well as the aspects that are being judged (akin to bills). Thus, PULSAR or future systems like it, open up new avenues for the systematic analysis of high-dimensional opinions and judgments at scale within existing ideal point models.