Multilingual Clustering of Streaming News

Multilingual Clustering of Streaming News
复制标题

DOI:
10.18653/v1/d18-1483
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Sebastião Miranda;Arturs Znotins;Shay B. Cohen;Guntis Barzdins
Sebastião Miranda;Arturs Znotins;Shay B. Cohen;Guntis Barzdins
中科院分区:
其他
文献类型:
--
作者:
Sebastião Miranda;Arturs Znotins;Shay B. Cohen;Guntis Barzdins

文献摘要

被引文献

相似文献

跨语言聚合新闻可以通过将多语言来源的文章聚合为连贯的故事来实现高效的媒体监控。在在线环境中这样做可以对大量新闻流进行可扩展的处理。为此,我们描述了一种新的方法聚类传入流的多语言文档成单语和跨语言集群。不同于典型的聚类方法,报告结果的数据集与一个小的和已知数量的标签,我们解决的问题,发现越来越多的集群标签在一个在线的方式,使用真实的新闻数据集在多种语言。在我们的公式中,单语集群将文档分组在一起,而跨语言集群将单语集群分组在一起,每个语言出现在流中。我们的方法易于实现,计算效率高,并在德语,英语和西班牙语的数据集上产生最先进的结果。
Clustering news across languages enables efficient media monitoring by aggregating articles from multilingual sources into coherent stories. Doing so in an online setting allows scalable processing of massive news streams. To this end, we describe a novel method for clustering an incoming stream of multilingual documents into monolingual and crosslingual clusters. Unlike typical clustering approaches that report results on datasets with a small and known number of labels, we tackle the problem of discovering an ever growing number of cluster labels in an online fashion, using real news datasets in multiple languages. In our formulation, the monolingual clusters group together documents while the crosslingual clusters group together monolingual clusters, one per language that appears in the stream. Our method is simple to implement, computationally efficient and produces state-of-the-art results on datasets in German, English and Spanish.