EAGER: Annotating and extracting detailed syntactic information from a 1.1-billion-word corpus
EAGER: Annotating and extracting detailed syntactic information from a 1.1-billion-word corpus
批准号:
2026850
负责人:
Beatrice Santorini
金额:
$29.84万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-08-15 至 2024-01-31
中文摘要
在过去的十年里,研究人员已经获得了非常大的英语文本语料库,事实证明这些语料库对语言科学具有相当大的价值。甚至在最近,自然语言处理的方法已经发展到了一个点,我们可以开始想象使用自动解析和未校正的语料库进行语言研究,迄今为止,这种语料库是使用人工校正的语料库进行的。正是这种新情况,私人投资机构希望利用这一新情况,在最近完成并可供研究的数字化早期英语在线丛书(EEBO)语料库的基础上,生成一个自动解析的10亿多个词的早期现代英语语料库。其目的是利用最近开发的自然语言处理的尖端方法,创建一个具有适合语言和计算研究的准确度的自动解析数据库。由此产生的资源将使迄今不可能进行的调查成为可能;具体地说,EEBO的解析版本中包含的信息将允许研究人员不仅调查单词的频率影响,而且调查更大的语法单位(短语和从句)的频率影响。除了它们固有的语言兴趣之外,这些研究的结果可能会导致发现更复杂的基于意义的属性以及这些属性是如何变化的,这对自然语言处理的研究应该是有价值的。投资促进机构在实现这一目标方面取得了进展,创建了EEBO语料库的第一个自动解析版本,并开始评估其准确性。一些特征,如从句否定的句法,已经在我们的能力范围内,但对于许多其他结构,大规模方法检索的精确度仍有待确定。由于EEBO甚至比最大的单个人工纠正的语料库大300多倍,预计比现在可用的更准确的解析版本将开始允许研究人员研究只在现有英语语料库中零星证实的现象,专注于历史变化的起点和终点,以迄今无法实现的准确性和可靠性调查许多不同类型的频率效应(包括已经提到的新的频率效应),并严格评估语言变化的数学模型。由于EEBO(1500-1700)所涵盖的英语阶段已经是公认的现代语言,EEBO的句法版本在某种程度上可以代表用于语言科学研究的现代英语语料库。因此,它应该是计算语言学中的应用程序的训练和测试平台,包括词性标记、句法分析、命名实体识别,最终是词汇化、词义消除歧义等。EEBO巨大的体裁多样性和可变的拼写,以及它与现代英语的适度距离,也将使其解析版本成为评估和改进这些应用程序的健壮性以及开发新的解析器评估度量标准的自然候选者,这些标准可以作为计算语言学的语言信息基准。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Over the past decade, very large text corpora of English have become available to researchers that turn out to be of considerable value for the language sciences. Even more recently, methods in natural language processing have advanced to a point where we can begin to imagine conducting linguistic research using automatically parsed and uncorrected corpora of the sort that has so far been conducted using human-corrected corpora. It is this new situation that the PIs wish to exploit by producing an automatically parsed billion-plus word corpus of early modern English based on the digitized Early English Books Online (EEBO) corpus that has recently been completed and made accessible to research. The aim is to create an automatically parsed database with a level of accuracy suitable for both linguistic and computational research, using the recently developed cutting-edge methods in natural language processing. The resulting resource will make possible investigations hitherto impossible; specifically, the information contained in a parsed version of EEBO will permit researchers to investigate frequency effects not just of words, but of larger grammatical units (phrases and clauses). In addition to their inherent linguistic interest, the results of such investigations may lead to the discovery of more sophisticated meaning-based properties and how these vary, which should be of value for research in natural language processing. The PIs have made progress on this goal, having created a first automatically parsed version of the EEBO corpus and begun to assess its accuracy. Some features like the syntax of clausal negation are already within our reach, but for many other structures, it remains to be determined how accurate retrieval with large-scale methods can be. Since EEBO is more than 300 times larger than even the largest individual human-corrected corpora, it is expected that a more accurately parsed version of it than the one now available will begin to allow researchers to study phenomena that are only sporadically attested in existing English corpora, to zero in on the very beginnings and ends of historical changes, to investigate many different types of frequency effects (including the novel ones already mentioned) with an accuracy and reliability not hitherto possible, and to rigorously evaluate mathematical models of language change. Because the stage of English covered by EEBO (1500-1700) is already recognizably the modern language, a parsed version of EEBO can to some extent stand proxy for a corpus of Present-Day English for research in the language sciences. As a result, it should be useful as a training and testing ground for applications in computational linguistics including part-of-speech tagging, parsing, named entity recognition, and eventually lemmatization, sense disambiguation, and others. EEBO’s great genre variety and variable orthography and its moderate distance from Present-Day English will also make a parsed version of it a natural candidate for assessing and improving the robustness of these applications and for developing novel parser evaluation metrics that can serve as linguistically informed benchmarks for computational linguistics.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Parsing Early Modern English for linguistic search
解析早期现代英语以进行语言搜索
DOI:
10.7275/twww-ef90
发表时间:
2022
期刊:
Proceedings of the Society for Computation in Linguistics
影响因子:
--
作者:
[Kulick, Seth, Ryant, Neville, Santorini, Beatrice]
通讯作者:
Santorini, Beatrice
Parsing "Early English Books Online" for linguistic search
解析“早期英语在线书籍”以进行语言搜索
DOI:
--
发表时间:
2023
期刊:
Proceedings of the Society for Computation in Linguistics
影响因子:
--
作者:
[Kulick, Seth, Ryant, Neville, Santorini, Beatrice]
通讯作者:
Santorini, Beatrice
Collaborative Research: A corpus of New York City English: Audio-aligned and parsed
-
批准号:1629348
-
项目类别:Standard Grant
-
资助金额:$8.01万
-
财政年份:2016
-
负责人:Beatrice Santorini
-
依托单位:
Collaborative Research: A syntactically annotated corpus of Appalachian English
-
批准号:1151630
-
项目类别:Standard Grant
-
资助金额:$4.65万
-
财政年份:2012
-
负责人:Beatrice Santorini
-
依托单位:
海外基金