Detecting Topics Popular in the Recent Past from a Closed Caption TV Corpus as a Categorized Chronicle Data

Detecting Topics Popular in the Recent Past from a Closed Caption TV Corpus as a Categorized Chronicle Data
复制标题

DOI:
10.5220/0005612103420349
复制
发表时间:
2015-11
期刊:
--
影响因子:
--
通讯作者:
H. Mochizuki;Kohji Shibano
H. Mochizuki;Kohji Shibano
中科院分区:
其他
文献类型:
--
作者:
H. Mochizuki;Kohji Shibano

文献摘要

相似文献

在本文中,我们提出了一种方法来提取我们感兴趣的主题在过去的28个月的过程中,从封闭式字幕电视语料库。每个电视节目都被分配了以下类型之一:戏剧,信息或小报风格的节目,音乐,电影,文化,新闻,综艺,福利或体育。本文的研究重点是信息/小报式的节目,戏剧和新闻。使用我们的方法,我们提取的bigrams,形成的一个女主角的签名短语的一部分,在一个流行的戏剧,以及最近的世界,国内,娱乐圈,等等新闻的一个英雄的名字。实验结果表明,本文提出的方法与LDA模型在主题检测方面具有相同的效果,同时,本文的字幕电视语料库具有丰富的文化和社会生活分类记录的潜在价值。
In this paper, we propose a method for extracting topics we were interested in over the course of the past 28 months from a closed-caption TV corpus. Each TV program is assigned one of the following genres: drama, informational or tabloid-style program, music, movie, culture, news, variety, welfare, or sport. We focus on informational/tabloid-style programs, dramas and news in this paper. Using our method, we extracted bigrams that formed part of the signature phrase of a heroine and the name of a hero in a popular drama, as well as recent world, domestic, showbiz, and so on news. Experimental evaluations show that our simple method is as useful as the LDA model for topic detection, and our closed-caption TV corpus has the potential value to act as a rich, categorized chronicle for our culture and social life.