Towards a Model for Discourse Marker Annotation in spoken French: From potential to feature-based discourse markers

Towards a Model for Discourse Marker Annotation in spoken French: From potential to feature-based discourse markers
复制标题

法语口语语篇标记注释模型:从潜力到基于特征的语篇标记

DOI:
--
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
Deniz Uygur
Deniz Uygur
中科院分区:
--
文献类型:
--
作者:
C. Bolly;Ludivine Crible;Liesbeth Degand;Deniz Uygur

文献摘要

被引文献

相似文献

从没有公认的话语标记语(DM)的封闭类和一些语言标记语可能会或可能不会被算作DM的定义(Schourup 1999:228)的共同观察出发,我们的目标是提出一种识别和注释自发法语口语中DM的经验方法(MDMA项目)。我们的建议的核心是,DM可以被描述为集群的功能,在特定的组合模式,允许区分DM的使用从其他用途。我们分三步进行:(i)使用非常广泛的话语标记语定义,即“指示听话人如何将其宿主话语融入话语的发展中心理模型,使该话语看起来最佳连贯”的项目(汉森2006:25)-三位分析师在一份800字的文字记录中确定了所有潜在的话语标记;(ii)从一个10,000字的语料库中提取所有类型;(iii)根据10多个特征(包括句法、语义、搭配和韵律特征)进行分析。我们的注释实验的假设是,对特定标记施加的分布约束的分析应该发现可靠的功能,用于识别和分类的DM。在分析的第一步中提取的潜在话语标记语是指那些在某种语境中可以实现话语标记功能的语言表达,即在以下“层次”或“领域”中使用的语言表达:“对话的顺序结构,话轮转换系统,言语管理,人际管理,主题结构和参与框架”(Fischer 2006:9)。例如,tu vois 'you see'被定义为潜在的DM,因为它可以出现在用于管理说话者和听话者之间关系的上下文中(1),尽管在其他上下文中它没有(2)。根据研讨会的目标,我们的奋进旨在建立(更)可靠的标准来分类DM,以及如何将它们与其他履行非命题功能的语言项目区分开来,例如语气词(Degand et al. 2013)或语用标记(Brinton 1996)。注释者之间的分歧也揭示了一些边缘性的情况,尽管通常可以检测到一些话语、语用、索引功能,但对于这些项目属于哪一类别似乎有些犹豫(例如例3-4中的c 'est ca 'that's it ',quand meme 'still')。在本演示文稿中,我们首先简要介绍了我们的方法选择和问题,然后揭示了编码器间不一致的问题示例,以及特征聚类统计分析的第一个结果。后者表明,在根据上下文确定模式描述的过程中,所审查的不同特征在其相关性、可靠性或有用性方面有一定的等级。
Starting from the common observation that there is no recognized closed class of discourse markers (DMs) and that a number of linguistic markers may or may not count as DMs according to the definitions at stake (Schourup 1999: 228), we aim to present an empirical method for the identification and annotation of DMs in spontaneous spoken French (MDMA project). Central to our proposal is that DMs may be described as clusters of features that, in specific patterns of combination, allow distinguishing DM use from other uses. We proceeded in three steps: (i) using a very broad definition of DMs – i.e. items that “provide instructions to the hearer on how to integrate their host utterance into a developing mental model of the discourse in such a way as to make that utterance appear optimally coherent” (Hansen 2006: 25) – three analysts identified all potential DMs in an 800 words transcript; (ii) all types found were then extracted from a balanced 10,000 words corpus; and (iii) analyzed according to more than 10 features (including syntactic, semantic, collocational, and prosodic features). The hypothesis underlying our annotation experiment is that the analysis of the distributional constraints imposed on specific markers should uncover reliable features for the identification and categorization of DMs. The potential DMs extracted at the first step of analysis refer to those linguistic expressions that can, in one context or another, fulfill a DM function, i.e. be used at either of the following “levels” or “domains”: “the sequential structure of the dialogue, the turn-taking system, speech management, interpersonal management, the topic structure, and participation frameworks” (Fischer 2006: 9). For example, tu vois ‘you see’ is defined as a potential DM because it can occur in contexts where it serves to manage the relationship between speaker and hearer (1), although in other contexts it does not (2). In line with the objectives of the workshop, our endeavor seeks to establish (more) reliable criteria for the categorization as DMs, and how to distinguish them from other linguistic items fulfilling a non-propositional function, such as modal particles (Degand et al. 2013) or pragmatic markers (Brinton 1996). Disagreement between annotators also reveals borderline cases where, although some discursive, pragmatic, indexical function is commonly detected, there seems to be some hesitation as to what category these items belong to (such as c'est ca ‘that’s it’, quand meme ‘still’ in examples 3-4). In this presentation, we first briefly go over our methodological choices and issues, and then uncover problematic examples of inter-coder disagreement, as well as the first results of the statistical analysis of clusters of features. The latter suggests that there is a certain hierarchy between the different features under scrutiny, regarding their relevance, reliability, or usefulness in the process of identifying DMs in context.