Building a Corpus for the Zaza–Gorani Language Family

Building a Corpus for the Zaza–Gorani Language Family
复制标题

为 Zaza-Gorani 语系构建语料库

DOI:
--
复制
发表时间:
2020
期刊:
Workshop on NLP for Similar Languages, Varieties and Dialects
影响因子:
--
通讯作者:
Sina Ahmadi
Sina Ahmadi
中科院分区:
--
文献类型:
--
作者:
Sina Ahmadi

文献摘要

被引文献

相似文献

由于当地社区和各种新闻网站的发展,加上网络的可访问性越来越高,沿着,一些濒临灭绝和资源较少的语言有机会在信息时代复兴。因此,Web被认为是一个巨大的资源,可以用来提取语言语料库,使研究人员能够进行各种研究的语言学和语言技术。扎扎-戈兰尼语(Zaza-Gorani language)是伊朗西北部诸语言的一个分支,目前尚无重要的语料库。出于创建一个,在本文中,我们提出了我们的努力,收集语料库中的Zazaki和Gorani语言包含超过1.6M和194 K单词标记,分别。这个语料库是公开的。
Thanks to the growth of local communities and various news websites along with the increasing accessibility of the Web, some of the endangered and less-resourced languages have a chance to revive in the information era. Therefore, the Web is considered a huge resource that can be used to extract language corpora which enable researchers to carry out various studies in linguistics and language technology. The Zaza–Gorani language family is a linguistic subgroup of the Northwestern Iranian languages for which there is no significant corpus available. Motivated to create one, in this paper we present our endeavour to collect a corpus in Zazaki and Gorani languages containing over 1.6M and 194k word tokens, respectively. This corpus is publicly available.