课题基金 / 基金详情

Reading concordances in the 21st century (RC21)

Reading concordances in the 21st century (RC21)
21世纪阅读索引(RC21)
批准号:
AH/X002047/1
负责人:
Michaela Mahlberg
金额:
$36.07万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

Michaela Mahlberg的其他基金

相似基金

相关文献

中文摘要
翻译
在当今的数字世界中,以电子形式交流的文本数量不断增加,对从文本中提取大规模意义的途径和方法的需求也越来越大。语料库语言学家长期以来一直在研究数字化文本,并已确定许多语言的特征是反复出现的模式。因此,单词“眼”可以与“奶油”、“测试”或“关闭”、“固定”等词一起出现。在语料库语言学中,这样的模式是通过语料库识别的,即以紧凑的格式显示一个词、短语或结构在一系列上下文中的多次出现的显示。然而,由于缺乏完善和明确的方法论,阅读协和的艺术尚未充分发挥其潜力。与此同时,语料库语言学家可用的语料库软件在算法方面几乎没有创新。该项目提出了一种21世纪阅读和谐词的创新方法。通过伯明翰大学和弗里德里希-亚历山大-埃尔兰根-纽伦堡大学的合作,我们将语料库语言学的理论工作优势与计算算法方面的专业知识结合起来,以开发一种系统的语料库阅读方法。我们将开发独立于工具的和声读取策略,并将开发相应的算法来半自动分析和谐线。我们将具体实施软件FlexiConc,以支持语料库语言学家组织和解释语料库的语料库。为了开发和测试我们的方法,我们将进行两个案例研究。第一个案例研究将重点研究小说中的肢体语言与非小说文本中的肢体语言。第二个案例研究将关注社交媒体上的政治论证,将其结果正式化为可用于自动论证挖掘的语料库查询。这两个案例研究都包括英语和德语之间的比较维度。因此,它们拓宽了到目前为止非常专注于英语语言的和谐阅读的方法。通过这些案例研究,我们将建立一种方法,不仅提供语料库语言学的创新,而且对大规模文本数据的分析具有更广泛的影响,同时仍保留人文视角。我们将把FlexiConc开发为开源软件,以便其他研究人员可以将其作为现成的工具使用,或者将其集成到现有的整合工具或他们自己的软件环境中。FlexiConc和我们的工具无关的一致性分析方法将超越语料库语言学,为数字人文和计算社会科学等学科提供创新的方法和算法。我们将通过各种形式提高对新可能性的认识,例如,通过项目博客,我们软件的用户可以分享他们的经验,并在由领先的国际专家组成的咨询委员会的帮助下。我们将在暑期学校和会议上举办培训课程,并在网上提供教育材料。
英文摘要
In today's digital world, the amount of text communicated in electronic form is ever-increasing and there is a growing need for approaches and methods to extract meanings from texts at scale. Corpus linguists have long been studying digitised texts and have established that much of language is characterised by recurring patterns. So the word 'eye' can appear together with words like 'cream' and 'test', or words like 'closed' and 'fixed'. In corpus linguistics, such patterns are identified with the help of concordances, i.e. displays that show many occurrences of a word, phrase or construction across a range of contexts in a compact format. However, lacking a well-established and clear-cut methodology, the art of reading concordances has not yet realised its full potential. At the same time, there has been very little innovation in algorithms in the concordance software packages available to corpus linguists. This project proposes an innovative approach to reading concordances in the 21st century. Through the collaboration between the University of Birmingham and Friedrich-Alexander-Universität Erlangen-Nürnberg we combine strengths in theoretical work in corpus linguistics with expertise in computational algorithms in order to develop a systematic methodology for reading concordances. We will develop tool-independent strategies for reading concordances and we will develop corresponding algorithms for the semi-automatic analysis of concordance lines. We will specifically implement the software FlexiConc to support the corpus linguist researcher in organising and interpreting concordances. To develop and test our approach, we will conduct two case studies. The first case study will focus on body language in fiction compared to non-fiction texts. The second case study will focus on political argumentation in social media, formalising its findings as corpus queries that can be used for automatic argumentation mining. Both case studies include a comparative dimension between English and German. Hence, they broaden out approaches to concordance reading which have been very focused on the English language so far. Through these case studies, we will establish an approach that not only provides innovation in corpus linguistics, but also has wider implications for the analysis of textual data at scale, while still retaining a humanities perspective. We will develop FlexiConc as open-source software, so that other researchers can use it as an off-the-shelf tool or integrate it into existing concordance tools or their own software environment. Both FlexiConc and our tool-independent approach to concordance analysis will have relevance beyond corpus linguistics, providing innovative approaches and algorithms for disciplines such as digital humanities and computational social science. We will raise awareness of the new possibilities in a variety of forms, for instance, through a project blog where users of our software can share their experience, and with the help of an advisory board of leading international experts. We will run training sessions at summer schools and conferences and make educational materials available online.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CLiC Dickens - characterisation in the representation of speech and body language from a corpus stylistic perspective.
  • 批准号:
    AH/P504634/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $7.55万
  • 财政年份:
    2017
  • 负责人:
    Michaela Mahlberg
  • 依托单位:
CLiC Dickens - characterisation in the representation of speech and body language from a corpus stylistic perspective.
  • 批准号:
    AH/K005146/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $25.54万
  • 财政年份:
    2013
  • 负责人:
    Michaela Mahlberg
  • 依托单位:
海外基金