CzeDLex – A Lexicon of Czech Discourse Connectives

CzeDLex – A Lexicon of Czech Discourse Connectives
复制标题

CzeDLex – 捷克语话语连接词词典

DOI:
10.1515/pralin-2017-0039
复制
发表时间:
2017
期刊:
The Prague Bulletin of Mathematical Linguistics
影响因子:
--
通讯作者:
Lucie Poláková
Lucie Poláková
中科院分区:
--
文献类型:
--
作者:
Jirí Mírovský;Pavlína Synková;Magdaléna Rysová;Lucie Poláková

文献摘要

被引文献

相似文献

摘要 CzeDLex 是一个新的捷克语篇章连接词电子词典,计划于今年年底出版。其数据格式和结构基于对类似现有资源的研究,并进行了调整以符合捷克句法传统和细节以及布拉格文本中语义话语关系注释的方法。在本文中,我们首先将词典置于相关资源的背景下,并讨论构建词典的理论方面——我们提出了数据结构选择和词典条目特征选择的论据,同时特别注意主要连接词(例如英语中的“because”、“therefore”)和次要连接词(例如“for this Reason”,this is the Reason Why)的一致且(尽可能)统一的编码。词典中嵌套条目所采用的主要原则是——除了连接词的词汇形式之外——由给定连接词表达的话语语义类型(意义),这使我们能够处理连接词的广泛形式变化,并且便于将 CzeDLex 与其他语言的词典互连。其次,我们介绍了基于布拉格标记语言的技术解决方案,该解决方案可以将词典有效地纳入布拉格树库家族中——它可以在树编辑器 TrEd 中直接打开和编辑,在 btred 中通过命令行进行处理,与其源语料库互连并在 PML 树查询引擎中查询。第三,我们描述了通过利用手动标注语篇关系的大型语料库——布拉格语篇树库2.0来获取词典数据的过程:我们详细阐述了自动提取部分、提取后检查和手动添加补充语言信息。
Abstract CzeDLex is a new electronic lexicon of Czech discourse connectives, planned for publication by the end of this year. Its data format and structure are based on a study of similar existing resources, and adjusted to comply with the Czech syntactic tradition and specifics and with the Prague approach to the annotation of semantic discourse relations in text. In the article, we first put the lexicon in context of related resources and discuss theoretical aspects of building the lexicon – we present arguments for our choice of the data structure and for selecting features of the lexicon entries, while special attention is paid to a consistent and (as far as possible) uniform encoding of both primary (such as in English because, therefore) and secondary connectives (e.g. for this reason, this is the reason why). The main principle adopted for nesting entries in the lexicon is – apart from the lexical form of the connective – a discoursesemantic type (sense) expressed by the given connective, which enables us to deal with a broad formal variability of connectives and is convenient for interlinking CzeDLex with lexicons in other languages. Second, we introduce the chosen technical solution based on the Prague Markup Language, which allows for an efficient incorporation of the lexicon into the family of Prague treebanks – it can be directly opened and edited in the tree editor TrEd, processed from the command line in btred, interlinked with its source corpus and queried in the PML Tree Query engine. Third, we describe the process of getting data for the lexicon by exploiting a large corpus manually annotated with discourse relations – the Prague Discourse Treebank 2.0: we elaborate on the automatic extraction part, post-extraction checks and manual addition of supplementary linguistic information.