Speech representation - A literary and linguistic corpus study
Speech representation - A literary and linguistic corpus study
批准号:
322751860
负责人:
Dr. Annelen Brunner
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2016
资助国家:
德国
项目状态:
已结题
起止时间:
2015-12-31 至 2020-12-31
中文摘要
该项目的目标是开发识别语音表征的自动方法,并根据其结果对语音表征模式和用法进行文学和语言学分析。该项目将展示数字人文方法如何在不放弃特定学科的研究兴趣的情况下,使文学和语言研究更紧密地融合在一起。方法论的重点在于开发(半)自动计算和语料库语言策略,并将其应用于小说、报纸和期刊语料库(时间焦点:1841-1918)。这些文本已经以数字格式提供,只需根据语料库进行改编。该项目可以建立在识别语音表示的现有原型的基础上,该原型使用基于规则的方法和机器学习,并在既定的编程框架(UIMA)中实施。在对部分语料库进行手动标注的基础上,对该原型进行了改进和增强。该项目完成后,将向科学界提供识别器以及手动注释的语料库文本。因此,该项目有助于开发用于自动标注大型语料库的自然语言处理工具。识别器将生成量化数据,首次允许在广泛的经验基础上对言语表征进行叙事学和语言学研究。这些数据让我们有机会探讨语言学和文学研究领域中的几个开放的研究问题。除了历史维度之外,不同文本类型(报纸文本-小说-杂志)的影响以及“高端”和“通俗”文学之间的区别尤其令人感兴趣。主要的理论研究问题是:1)直接表征、自由间接表征、间接表征和转述表征这四种言语表征的历时和语篇依存性发展;2)从词汇和结构的角度对查询公式的历时和语篇依存性发展;3)以言语表征动词为例研究语言变化机制。通过对这些问题的研究,该项目有助于叙事学、语篇类型/体裁研究和词汇论元结构的理论建设。
英文摘要
Goals of the project are developing automatic methods for the recognition of speech representation and, on the basis of their results, conducting literary and linguistic analyses of speech representation patterns and usage. The project will show how methods of Digital Humanities lead to a stronger convergence of literary and linguistic studies, without abandoning research interests specific to the disciplines. Methodological focus lies on developing (semi-)automatic computational and corpus linguistic strategies which are applied to a corpus of novels, newspapers and periodicals (temporal focus: 1841- 1918). These texts are already available in digital format and only need to be adapted for the corpus. The project can build upon an existing prototype for recognizing speech representation that uses rule-based methods and machine learning and is implemented in an established programming framework (UIMA). Based on a manual annotation of part of the corpus, this prototype will be improved and enhanced. Upon completion of the project the recognizer as well as the manually annotated corpus texts will be made available to the scientific community. Thus, the project contributes to the development of NLP tools for the automatic annotation of large corpora. The recognizer will generate quantitative data that allows for the first time a narratological and linguistic study of speech representation on a broad empirical basis. This data gives us a chance to approach several open research questions from the fields of linguistics and literary studies. Apart from the historical dimension, the influence of different text types (newspaper texts - fiction - magazines) and the distinction between 'high' and 'popular' literature are of particular interest. Key theoretical research questions are: 1) The diachronic and text-type dependent development of the four types of speech representation - direct, free indirect, indirect and reported representation; 2) the diachronic and text-type dependent development of inquit formula from a lexical and structural perspective; 3) speech representation verbs as an example for mechanisms of linguistic change. By working on those questions, the project contributes to theory construction in narratology, text type/genre studies and studies of lexical argument structures.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.5281/zenodo.4622037
发表时间:
2019
期刊:
影响因子:
--
作者:
[Brunner, Annelen , Weimer, Lukas , Ngoc Duyen Tanja , Engelberg, Stefan , Jannidis]
通讯作者:
Jannidis
国内基金
海外基金
稀疏表示及其在盲源分离中的应用研究
-
批准号:61104053
-
项目类别:青年科学基金项目
-
资助金额:23.0万元
-
批准年份:2011
-
负责人:杨祖元
-
依托单位:
约化群GL(n, F)的表示--F是非阿基米德局部域
-
批准号:10701034
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2007
-
负责人:覃瑜君
-
依托单位:
信号盲处理的稀疏表示方法
-
批准号:60475004
-
项目类别:面上项目
-
资助金额:23.0万元
-
批准年份:2004
-
负责人:李远清
-
依托单位: