A Proposal of a Keyword Extraction System for Detecting Social Issues

A Proposal of a Keyword Extraction System for Detecting Social Issues
复制标题

用于检测社会问题的关键词提取系统的提案

DOI:
10.13088/jiis.2013.19.3.001
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Mijung Kang
Mijung Kang
中科院分区:
--
文献类型:
--
作者:
Dami Jeong;Jaeseok Kim;Gi;Jong;Byung;Mijung Kang

文献摘要

被引文献

相似文献

为了发现失业、经济危机、社会福利等现代社会急需解决的重大社会问题,现有的研究方法通常是通过在线或离线调查收集专业专家和学者的意见。然而,这种方法似乎并不有效,从时间到时间。通常,由于费用问题,很少收集大量的调查答复。在某些情况下,也很难找到处理具体社会问题的专业人员。因此,样本集通常很小,可能会有一些偏差。此外,对于一个社会问题,由于每个专家都有自己的主观观点和不同的背景,几个专家可能会得出完全不同的结论。在这种情况下,很难弄清楚当前的社会问题是什么,哪些社会问题是真正重要的。为了克服目前的方法的缺点,在本文中,我们开发了一个原型系统,半自动检测社会问题的关键字,代表社会问题和问题,从约130万的新闻文章,约10个主要国内出版社在韩国,从2009年6月至2012年7月。我们所提出的系统包括(1)从所收集的新闻文章中收集和提取文本,(2)仅识别与社会问题相关的新闻文章,(3)分析韩语句子的词汇项,(4)基于概率主题建模随着时间的推移找到一组关于社会关键词的主题,(5)将相关段落与给定主题匹配,以及(6)可视化社会关键字以便于理解。特别是,我们提出了一种新的匹配算法依赖于生成模型。我们提出的匹配算法的目标是最好的匹配段落的每个主题。从技术上讲,使用潜在狄利克雷分配(LDA)等主题模型,我们可以获得一组主题,每个主题都有相关的术语及其概率值。在我们的问题中,给定一组文本文档(例如,新闻文章),LDA显示一组主题聚类,然后每个主题聚类由人类注释者标记,其中每个主题标签代表一个社交关键字。例如,假设存在主题(例如,专题1
To discover significant social issues such as unemployment, economy crisis, social welfare etc. that are urgent issues to be solved in a modern society, in the existing approach, researchers usually collect opinions from professional experts and scholars through either online or offline surveys. However, such a method does not seem to be effective from time to time. As usual, due to the problem of expense, a large number of survey replies are seldom gathered. In some cases, it is also hard to find out professional persons dealing with specific social issues. Thus, the sample set is often small and may have some bias. Furthermore, regarding a social issue, several experts may make totally different conclusions because each expert has his subjective point of view and different background. In this case, it is considerably hard to figure out what current social issues are and which social issues are really important. To surmount the shortcomings of the current approach, in this paper, we develop a prototype system that semi-automatically detects social issue keywords representing social issues and problems from about 1.3 million news articles issued by about 10 major domestic presses in Korea from June 2009 until July 2012. Our proposed system consists of (1) collecting and extracting texts from the collected news articles, (2) identifying only news articles related to social issues, (3) analyzing the lexical items of Korean sentences, (4) finding a set of topics regarding social keywords over time based on probabilistic topic modeling, (5) matching relevant paragraphs to a given topic, and (6) visualizing social keywords for easy understanding. In particular, we propose a novel matching algorithm relying on generative models. The goal of our proposed matching algorithm is to best match paragraphs to each topic. Technically, using a topic model such as Latent Dirichlet Allocation (LDA), we can obtain a set of topics, each of which has relevant terms and their probability values. In our problem, given a set of text documents (e.g., news articles), LDA shows a set of topic clusters, and then each topic cluster is labeled by human annotators, where each topic label stands for a social keyword. For example, suppose there is a topic (e.g., Topic1