CoNLL 2017 Shared Task - Automatically Annotated Raw Texts and Word Embeddings

CoNLL 2017 Shared Task - Automatically Annotated Raw Texts and Word Embeddings
复制标题

CoNLL 2017 共享任务 - 自动注释的原始文本和词嵌入

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Daniel Zeman
Daniel Zeman
中科院分区:
--
文献类型:
--
作者:
Filip Ginter;Jan Hajic;Juhani Luotolahti;Milan Straka;Daniel Zeman

文献摘要

被引文献

相似文献

由UDPipe (http://ufal.mff.cuni)生成的45种语言的原始文本的自动分割、标记、形态学和句法注释。Cz /udpipe),以及由word2vec (https://code.google.com/archive/p/word2vec/)从小写文本计算的维度为100的词嵌入。
Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts in 45 languages, generated by UDPipe (http://ufal.mff.cuni.cz/udpipe), together with word embeddings of dimension 100 computed from lowercased texts by word2vec (https://code.google.com/archive/p/word2vec/). For each language, automatic annotations in CoNLL-U format are provided in a separate archive. The word embeddings for all languages are distributed in one archive. Note that the CC BY-SA-NC 4.0 license applies to the automatically generated annotations and word embeddings, not to the underlying data, which may have different license and impose additional restrictions. Update 2018-09-03 =============== Added data in the 4 “surprise languages” from the 2017 ST: Buryat, Kurmanji, North Sami and Upper Sorbian. This has been promised before, during CoNLL-ST 2018 we gave the participants a link to this record saying the data was here. It wasn't, sorry. But now it is.