课题基金 / 基金详情

Building a parsed historical corpus to investigate word-order change and variation

Building a parsed historical corpus to investigate word-order change and variation
构建经过解析的历史语料库来研究词序变化和变异
批准号:
2314522
负责人:
Christopher Sapp
金额:
$45.8万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
活着的语言随着时间的推移在许多方面都会发生变化,不仅包括词汇和发音,还包括句子结构。历史语言学致力于记录这些变化并寻求对它们的解释。句子结构的变化通常发生在一段较长的时间内,包括在表达基本概念的各种语法模式之间存在差异的时期。在引入录音之前,这些变化的唯一证据是书面文件。然而,要从书面文件中收集足够的证据,对给定语言历史上的语法模式的变异和变化进行严格的科学调查,需要检查一个大型的、经过分析的语料库--一个分为句子、从句和短语的文本的集合。这个项目建立了一个单一语言的经过分析的电子语料库,涵盖了多个世纪、地理区域和文本体裁。这使得人们能够调查这种语言历史上的语法变化和变异,并与相关语言中类似的发展进行比较。该语料库可供任何研究人员公开使用,大学和高中的宣传活动促进了公众对使用科学和技术来探索语言结构问题的认识。语料库的开发还有助于培养下一代语言学研究人员,包括博士后研究员、研究生和本科生。该项目建立了一个140万字的句法分析电子语料库,包括165篇跨越1050-1950年和十个方言地区的文本,以及一系列的文本体裁。这需要对基于先前句法分析语料库的现有标注方案进行实质性扩展,以适应更广泛的句法现象,同时还保持标注方案与其他语言的少数句法分析历史语料库中使用的标注方案尽可能地可比。这个项目涉及文本的手动注释,纠正自动词性分析中出现的错误,消除许多句子的歧义,并交叉检查句法注释的准确性。由此产生的带注释的语料库填补了经分析的语料库集合中的空白,世界各地的研究人员可以免费使用该语料库以及有关语料库使用的文档。这个项目产生的经验数据为研究类型学变化的机制和在密切相关的方言中的传播提供了信息。语料库不仅可以用来研究句法领域的现象,也可以用来研究句法和其他语法成分之间的界面现象。鉴于语料库中的文本范围广泛,这些现象可以被同步地、历时地、社会语言学地研究,并与其他语言进行比较。这一奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Living languages change over time in a number of areas, including not only vocabulary and pronunciation, but also sentence structure. Historical linguistics is concerned with documenting these changes and seeking explanations for them. Changes in sentence structure often occur over an extended period of time, including a period in which there is variation between various grammatical patterns for expressing a basic notion. The only evidence for these changes before the introduction of sound recording consists of written documents. However, gathering sufficient evidence from written documents for a rigorous scientific investigation of variation and change in grammatical patterns in the history of a given language requires the examination of a large, parsed corpus — a collection of texts that is divided into sentences, clauses, and phrases. This project builds a parsed electronic corpus of a single language, covering multiple centuries, geographical areas, and text genres. This allows for the investigation of grammatical change and variation in the history of the language as well as comparison with similar developments in related languages. The corpus is publicly available for any researcher to use, and outreach to universities and high schools promotes public awareness of the use of science and technology to explore questions about the structure of language. The development of the corpus also contributes to the training of the next generation of researchers in linguistics including a postdoctoral researcher, graduate students, and undergraduate students. This project builds a 1.4-million-word syntactically parsed electronic corpus including 165 texts spanning the years 1050-1950 and ten dialectal regions, and a range of text genres. This requires substantial extension of existing annotation schemes based on previous syntactically parsed corpora to accommodate a broader range of syntactic phenomena, while also keeping the annotation scheme as comparable as possible with those used in the handful of syntactically parsed historical corpora of other languages. This project involves manual annotation of texts, correcting errors that arise in automatic part-of-speech parsing, disambiguation of many sentences, and cross-checking for accuracy of syntactic annotations. The resulting annotated corpus fills a gap among the set of parsed corpora the world's languages and is available free of charge to researchers around the world, together with documentation on the use of the corpus. The empirical data generated by this project informs research on the mechanisms and spread of typological change over time across closely related dialects. The corpus can be used to investigate phenomena not only in the domain of syntax, but also in the interfaces between syntax and other components of grammar. Given the broad range of texts in the corpus, these phenomena can be examined synchronically, diachronically, sociolinguistically, and in comparison with other languages.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金