Turkish Treebanking: Unifying and Constructing Efforts
Turkish Treebanking: Unifying and Constructing Efforts
复制标题
土耳其树库:统一和建设努力
DOI:
10.18653/v1/w19-4019
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Arzucan Özgür
中科院分区:
文献类型:
--
作者:
Utku Türk;Furkan Atmaca;S. Özates;Abdullatif Köksal;Balkiz Öztürk Basaran;Tunga Güngör;Arzucan Özgür
In this paper, we present the current version of two different treebanks, the re-annotation of the Turkish PUD Treebank and the first annotation of the Turkish National Corpus Universal Dependency (henceforth TNC-UD). The annotation of both treebanks, the Turkish PUD Treebank and TNC-UD, was carried out based on the decisions concerning linguistic adequacy of re-annotation of the Turkish IMST-UD Treebank (Türk et. al., forthcoming). Both of the treebanks were annotated with the same annotation process and morphological and syntactic analyses. The TNC-UD is planned to have 10,000 sentences. In this paper, we will present the first 500 sentences along with the annotation PUD Treebank. Moreover, this paper also offers the parsing results of a graph-based neural parser on the previous and re-annotated PUD, as well as the TNC-UD. In light of the comparisons, even though we observe a slight decrease in the attachment scores of the Turkish PUD treebank, we demonstrate that the annotation of the TNC-UD improves the parsing accuracy of Turkish. In addition to the treebanks, we have also constructed a custom annotation software with advanced filtering and morphological editing options. Both the treebanks, including a full edit-history and the annotation guidelines, and the custom software are publicly available under an open license online.