Semantic Annotation and Mark Up for Enhancing Lexical Searches (SAMUELS)
Semantic Annotation and Mark Up for Enhancing Lexical Searches (SAMUELS)
批准号:
AH/L010062/1
负责人:
Marc Alexander
金额:
$51.78万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2014
资助国家:
英国
项目状态:
已结题
起止时间:
2014 至 --
中文摘要
随着人文学科数据集变得越来越大,研究人员迫切需要更复杂的分析技术。在文本数据集的大数据研究中,最重要的问题是,我们搜索、聚合和分析文本数据集的主要方法不是依赖于概念或含义,而是依赖于词的形式。这些形式对于它们所指代的意思来说是不完美的、难以回避的代理,英语中60%的单词形式指代不止一个意思,有些单词形式指代接近200个意思,使用单词形式进行搜索时出现的不相关“噪音”随着搜索文本的大小而增加。在大数据环境中,这个问题阻碍了研究,使任何详细的分析都变得完全难以处理,需要大量的人工干预。在这个项目中,我们将提供一个系统,用于自动标注文本中的单词及其精确含义,从而实现我们处理大型文本数据的方式的逐步改变。该系统基于无与伦比的《英语历史同义词典》,该词典包含了英语历史上79.7万个单词,按23.6万个层次分类,并附有每个单词的已知使用日期。注释软件将获取文本,并为文本中包含的每个单词提供XML注释,给出单词含义的Historical Thesaurus类别代码。该系统将使用一系列最先进的计算技术和新的上下文相关方法自动消除词义歧义,这些方法是由Thesaurus的年代代码和其独特的详细和细粒度的层次结构解锁的。以这种方式标记的文本数据可以被准确地搜索和精确地调查,任何结果也可以在一系列精度级别上进行汇总,而不需要人工干预。该项目的一个主要部分也是开发用于处理语义聚合和消歧数据的新技术。项目合作伙伴将对包括Hansard语料库(包含超过23亿单词的文本)、牛津英语语料库(世界上最大的现代英语分层语料库)和EEBO-TCP语料库(包含4万本早期现代书籍)在内的资源进行研究。作为我们改变处理这种规模数据的方式的一部分,我们将挖掘这些文本集合中频繁出现的或统计上不寻常的概念,将利用我们在大型数据集中搜索由歧义词形式实现的术语的能力(例如“union”在工业关系的特定上下文中,而不是这个词的其他33种可能的含义中的任何一种),并将从远读的角度审视整个数据,以寻找随着时间的推移,意义变化的显著或重要模式。这些基于标记数据的研究项目也将推动我们使用这些数据的工具的发展,英国和国外的研究团队提供了一系列不同的数据需求,确保在项目开发中满足各种需求和用例。通过这种方式,我们致力于在项目的生命周期内使用语义标记的数据产生一系列引人注目的、富有成效的和实用的研究成果,以证明我们的方法的价值,并帮助确保项目的工作得到尽可能广泛的利用和利用。通过做所有这些,我们将启用新的和变革性的技术来探索、搜索和调查大型人文数据集中的大规模文化、文学、历史和语言现象;通过这个项目,将有可能把意义——而不是文字形式——置于数字人文研究文本的核心。
英文摘要
As humanities datasets get ever larger, researchers have a pressing need for more sophisticated techniques of analysis. The most significant issue in big data research into textual datasets is that our primary methodology for searching, aggregating and analysing them relies not on concepts or meanings, but rather on word forms. These forms are imperfect and evasive proxies for the meanings they refer to, and with 60% of word forms in English referring to more than one meaning, and some word forms referring to close to two hundred meanings, the irrelevant "noise" which appears when searching using word forms grows with the size of the texts being searched.In big data contexts, this problem cripples research, making any sort of detailed analysis entirely intractable and requiring impossible amounts of manual intervention. In this project, we will deliver a system for automatically annotating words in texts with their precise meanings, enabling a step-change in the way we deal with large textual data. The system is based around the unparalleled Historical Thesaurus of English, which contains 797,000 words from across the history of English arranged into 236,000 hierarchical categories of meanings alongside each word's dates of known use. The annotation software will take a text and provide for each word it contains an XML annotation giving the word meaning's Historical Thesaurus category code. The system will automatically disambiguate word meanings using a range of state-of-the-art computational techniques alongside new context-dependent methods unlocked by the Thesaurus's dating codes and its uniquely detailed and fine-grained hierarchical structure.Textual data tagged in this way can then be accurately searched and precisely investigated, with any results also able to be aggregated at a range of levels of precision, without the need for manual intervention. A major part of the project is also the development of new techniques for working with semantically-aggregated and disambiguated data. Project partners will conduct research on resources including the Hansard Corpus, consisting of over 2.3 billion words of text, the Oxford English Corpus, the world's largest stratified corpus of modern English, and the EEBO-TCP corpus of 40,000 early modern books. As part of our work on changing the nature of how we deal with data on this scale, we will mine these text collections for frequently-occurring or statistically unusual concepts, will take advantage of our ability to search large datasets for terms realised by ambiguous word forms (such as "union" in the particular context of industrial relations rather than any of the other 33 possible meanings of this word), and will examine the data as a whole from a distant-reading perspective in order to look for striking or significant patterns of meaning changes across time.These research projects based on tagged data will also drive the development of our tools for using this data, with teams of researchers across the UK and abroad providing a range of different demands on the data, ensuring a variety of needs and use-cases are catered for in the development of the project. In this way, we are committed to producing a set of compelling, fruitful, and practical research outcomes using semantically-tagged data during the lifetime of the project, in order to demonstrate the value of our approach and to help ensure the work of the project is as widely utilised and exploited as possible.By doing all of this, we will enable new and transformative techniques of exploring, searching and investigating large-scale cultural, literary, historical and linguistic phenomena in big humanities datasets; through this project, it will be possible to place meaning - rather than word forms - at the heart of digital humanities research into text.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Impression management in the Early Modern English courtroom
早期现代英国法庭中的印象管理
DOI:
10.1075/jhp.00019.arc
发表时间:
2018
期刊:
Journal of Historical Pragmatics
影响因子:
0.8
作者:
[Archer D]
通讯作者:
Archer D
"In barbarous times and in uncivilized countries" Two centuries of the evolving uncivil in the Hansard Corpus
“在野蛮时代和不文明国家”《国会议事录》语料库中两个世纪以来不断演变的不文明行为
DOI:
10.1075/ijcl.22016.ale
发表时间:
2022
期刊:
International Journal of Corpus Linguistics
影响因子:
1
作者:
[Alexander M]
通讯作者:
Alexander M
Mapping Hansard Impression Management Strategies through Time and Space
通过时间和空间映射国会议事印象管理策略
DOI:
10.1080/00393274.2017.1370981
发表时间:
2017
期刊:
Studia Neophilologica
影响因子:
0.4
作者:
[Archer D]
通讯作者:
Archer D
DOI:
10.4324/9781003031758
发表时间:
2020-04
期刊:
影响因子:
--
作者:
[S. Adolphs;Dawn Knight]
通讯作者:
S. Adolphs;Dawn Knight
Metaphor, Popular Science, and Semantic Tagging: Distant reading with the Historical Thesaurus of English
隐喻、科普和语义标签:利用英语历史词库进行远读
DOI:
10.1093/llc/fqv045
发表时间:
2015
期刊:
Digital Scholarship in the Humanities
影响因子:
0.8
作者:
[Alexander M]
通讯作者:
Alexander M
共 10 条
The formulation and management of social problems in service provision
-
批准号:ES/T008172/1
-
项目类别:Fellowship
-
资助金额:$11.68万
-
财政年份:2019
-
负责人:Marc Alexander
-
依托单位:
海外基金