Edinburgh Research Explorer Developing an automatic part-of-speech tagger for Scottish Gaelic

Edinburgh Research Explorer Developing an automatic part-of-speech tagger for Scottish Gaelic
复制标题

爱丁堡研究探索者为苏格兰盖尔语开发自动词性标注器

DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
W. Lamb
W. Lamb
中科院分区:
--
文献类型:
--
作者:
S. Danso;W. Lamb

文献摘要

被引文献

相似文献

本文描述了一个正在进行的项目,旨在为苏格兰盖尔语开发第一个自动PoS标记器。为爱尔兰语调整假释标记集,我们手动重新标记了一个已有的86k苏格兰盖尔语标记语料库。一个由13.5k个令牌组成的双重验证子集用于实例化8个统计标记器,并通过随机分配的保留样本验证它们的准确性。使用Brill双字标注器,准确率达到76.6%。我们概述了该项目的方法、中期结果和未来方向
This paper describes an on-going project that seeks to develop the first automatic PoS tagger for Scottish Gaelic. Adapting the PAROLE tagset for Irish, we manually re-tagged a pre-existing 86k token corpus of Scottish Gaelic. A double-verified subset of 13.5k tokens was used to instantiate eight statistical taggers and verify their accuracy, via a randomly assigned hold-out sample. An accuracy level of 76.6% was achieved using a Brill bigram tagger. We provide an overview of the project’s methodology, interim results and future directions