A Finnish news corpus for named entity recognition

A Finnish news corpus for named entity recognition
复制标题

DOI:
10.1007/s10579-019-09471-7
复制
发表时间:
2020-03-01
影响因子:
2.7
通讯作者:
Linden, Krister
Linden, Krister
中科院分区:
计算机科学4区
文献类型:
--
作者:
Ruokolainen, Teemu;Kauppinen, Pekka;Linden, Krister

文献摘要

被引文献

相似文献

我们提供了一个芬兰语新闻文章的语料库,其中包含手动准备的命名实体注释。语料库由953篇文章(193,742个单词令牌)和6个命名实体类(组织、位置、人员、产品、事件和日期)组成。这些文章摘自芬兰在线科技新闻来源Digitoday的档案。该语料库可用于研究目的。我们在两个域内和域外测试集上使用基于规则和两个深度学习系统在语料库上进行了基线实验。
We present a corpus of Finnish news articles with a manually prepared named entity annotation. The corpus consists of 953 articles (193,742 word tokens) with six named entity classes (organization, location, person, product, event, and date). The articles are extracted from the archives of Digitoday, a Finnish online technology news source. The corpus is available for research purposes. We present baseline experiments on the corpus using a rule-based and two deep learning systems on two, in-domain and out-of-domain, test sets.