课题基金 / 基金详情

Building a parsed historical corpus to investigate word-order change and variation

Building a parsed historical corpus to investigate word-order change and variation
构建经过解析的历史语料库来研究词序变化和变异
批准号:
2314522
负责人:
Christopher Sapp
金额:
$45.8万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
随着时间的推移,现存的语言在许多方面都在变化,不仅包括词汇和发音,还包括句子结构。历史语言学关注的是记录这些变化并为其寻找解释。句子结构的变化通常发生在一段较长的时间内,包括表达基本概念的各种语法模式之间存在变化的时期。在引入录音之前,这些变化的唯一证据是书面文件。然而,从书面文件中收集足够的证据,以便对特定语言历史上语法模式的变化和变化进行严格的科学调查,需要检查一个大的、已解析的语料库——一个分为句子、分句和短语的文本集合。这个项目建立了一个单一语言的解析电子语料库,涵盖多个世纪、地理区域和文本类型。这样就可以研究语言历史上的语法变化和变异,并与相关语言的类似发展进行比较。该语料库可供任何研究人员公开使用,并向大学和高中推广使用科学技术来探索语言结构问题的公众意识。语料库的开发也有助于培养下一代语言学研究人员,包括博士后研究员、研究生和本科生。该项目建立了一个140万字的语法解析电子语料库,包括165个文本,跨越1050-1950年,10个方言区域,以及一系列文本体裁。这就需要基于以前的语法解析语料库对现有注释方案进行大量扩展,以适应更广泛的语法现象,同时还要使注释方案尽可能与少数其他语言的语法解析历史语料库中使用的注释方案具有可比性。本项目涉及文本的手工注释,纠正自动词性分析中出现的错误,消除许多句子的歧义,并交叉检查句法注释的准确性。由此产生的注释语料库填补了世界语言解析语料库之间的空白,并免费提供给世界各地的研究人员,以及使用语料库的文档。该项目产生的经验数据为研究密切相关的方言类型变化的机制和传播提供了信息。语料库不仅可以用来研究语法领域的现象,还可以用来研究语法与其他语法组成部分之间的接口。鉴于语料库中文本的广泛范围,这些现象可以从共时性、历时性、社会语言学以及与其他语言的比较中进行检查。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Living languages change over time in a number of areas, including not only vocabulary and pronunciation, but also sentence structure. Historical linguistics is concerned with documenting these changes and seeking explanations for them. Changes in sentence structure often occur over an extended period of time, including a period in which there is variation between various grammatical patterns for expressing a basic notion. The only evidence for these changes before the introduction of sound recording consists of written documents. However, gathering sufficient evidence from written documents for a rigorous scientific investigation of variation and change in grammatical patterns in the history of a given language requires the examination of a large, parsed corpus — a collection of texts that is divided into sentences, clauses, and phrases. This project builds a parsed electronic corpus of a single language, covering multiple centuries, geographical areas, and text genres. This allows for the investigation of grammatical change and variation in the history of the language as well as comparison with similar developments in related languages. The corpus is publicly available for any researcher to use, and outreach to universities and high schools promotes public awareness of the use of science and technology to explore questions about the structure of language. The development of the corpus also contributes to the training of the next generation of researchers in linguistics including a postdoctoral researcher, graduate students, and undergraduate students. This project builds a 1.4-million-word syntactically parsed electronic corpus including 165 texts spanning the years 1050-1950 and ten dialectal regions, and a range of text genres. This requires substantial extension of existing annotation schemes based on previous syntactically parsed corpora to accommodate a broader range of syntactic phenomena, while also keeping the annotation scheme as comparable as possible with those used in the handful of syntactically parsed historical corpora of other languages. This project involves manual annotation of texts, correcting errors that arise in automatic part-of-speech parsing, disambiguation of many sentences, and cross-checking for accuracy of syntactic annotations. The resulting annotated corpus fills a gap among the set of parsed corpora the world's languages and is available free of charge to researchers around the world, together with documentation on the use of the corpus. The empirical data generated by this project informs research on the mechanisms and spread of typological change over time across closely related dialects. The corpus can be used to investigate phenomena not only in the domain of syntax, but also in the interfaces between syntax and other components of grammar. Given the broad range of texts in the corpus, these phenomena can be examined synchronically, diachronically, sociolinguistically, and in comparison with other languages.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金