课题基金 / 基金详情

Corpws Cenedlaethol Cymraeg Cyfoes (The National Corpus of Contemporary Welsh): A community driven approach to linguistic corpus construction

Corpws Cenedlaethol Cymraeg Cyfoes (The National Corpus of Contemporary Welsh): A community driven approach to linguistic corpus construction
Corpws Cenedlaethol Cymraeg Cyfoes(当代威尔士语国家语料库):社区驱动的语言语料库建设方法
批准号:
ES/M011348/1
负责人:
Dawn Knight
金额:
$183.6万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2016
资助国家:
英国
项目状态:
已结题
起止时间:
2016 至 --

项目摘要

项目成果

Dawn Knight的其他基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This project will create a major corpus of Welsh language: CorCenCC (Corpws Cenedlaethol Cymraeg Cyfoes: National Corpus of Contemporary Welsh). A corpus is a principled collection of language data sampled from real-life contexts, presented as a searchable database. This will be the first corpus to represent spoken, written and electronically-mediated Welsh, and the first in any language with a functional design informed, from the outset, by representatives of all anticipated academic and community user groups. CorCenCC will provide societal, economic and academic benefits by:- Facilitating uses of Welsh in public, commercial, educational and governmental settings.- Redefining the scope, relevance and design infrastructure of corpus development methodology.A corpus allows users to identify and explore language as it is actually used, rather than relying on intuition or prescriptive accounts of how it 'should' be used. This evidence-based approach is used by academic researchers, lexicographers, teachers, language learners, assessors, resource developers, policy makers, publishers, translators and others, and is essential to the development of technologies such as predictive text production, word processing tools, machine translation, voice recognition and web search tools. Welsh has had no comprehensive corpus facility able to meet these requirements.CorCenCC will capitalise on extensive community interest in sustaining and 'growing' Welsh, using the novel integration of crowdsourcing, a powerful data collection method which has the potential to revolutionize corpus construction. Recruited through social and broadcast media, roadshows and existing networks, Welsh speakers will record and upload their own data via a mobile app, and even contribute to data coding. This approach promises representative language across genres, language varieties (regional and social) and contexts. Traditional, data collection will supplement the crowdsourcing, ensuring a representative balance of data as specified in the project targets.Preliminary engagement with stakeholders (including a briefing event at the Senedd) generated collaboration from the Welsh Government, Welsh Language Commissioner, Welsh Joint Education Committee, Welsh for Adults, BBC, Gwasg y Lolfa press, and University of Wales Dictionary; all have identified current needs which CorCenCC can meet, and all will be represented in the project advisory group, so the corpus design is user-informed throughout. A language corpus able to inform delivery of Welsh has been called for by e.g. National Foundation for Educational Research (2008:48) and Welsh Government (2013:27,71). CorCenCC, with its integrated pedagogical toolkit, will impact significantly on Welsh language teaching practice, enabling data-driven, inductive learning and assessment.CorCenCC will be open-source and publicly accessible, with user interfaces for specific groups. It will enable, for example, community users to investigate dialect variation or idiosyncrasies of their own language use; professional users to profile texts for readability or develop digital language tools; language learners learn from real life models of Welsh; and researchers to investigate patterns of language use and change. In order to ensure that CorCenCC remains a sustainable, permanent and user-oriented record of language, an in-built facility will allow data to be added and moderated beyond the life of the project. The project team comprises experts in corpus linguistics, Welsh, and language pedagogy and assessment, who specialise in the application of linguistic tools to real world issues. Working with an advisory body of stakeholder representatives, they are optimally placed to meet the project aims: creating a permanent, sustainable and fit-for-purpose record of the living language, and pioneering an approach to content generation and user-driven applications that will provide a model for future corpus creation.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
CorCenCC (Corpws Cenedlaethol Cymraeg Cyfoes - National Corpus of Contemporary Welsh): A demonstration
CorCenCC(Corpws Cenedlaethol Cymraeg Cyfoes - 当代威尔士语国家语料库):演示
DOI: --
发表时间: 2018
期刊:
影响因子: --
作者: [Knight D]
通讯作者: Knight D
DOI: 10.3390/app11156896
发表时间: 2021-07
期刊: Applied Sciences
影响因子: --
作者: [P. Corcoran;Geraint I. Palmer;Laura Arman;Dawn Knight;Irena Spasic]
通讯作者: P. Corcoran;Geraint I. Palmer;Laura Arman;Dawn Knight;Irena Spasic
DOI: 10.18653/v1/w19-4332
发表时间: 2019-08
期刊:
影响因子: --
作者: [I. Ezeani;S. Piao;Steven Neale;Paul Rayson;Dawn Knight]
通讯作者: I. Ezeani;S. Piao;Steven Neale;Paul Rayson;Dawn Knight
Creating pedagogical wordlists: a comparison of thematic and corpus approaches
创建教学词汇表:主题方法和语料库方法的比较
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者: [Fitzpatrick T]
通讯作者: Fitzpatrick T
9
    FreeTxt: supporting bilingual free-text survey and questionnaire data analysis
    • 批准号:
      AH/W004844/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $10.28万
    • 财政年份:
      2022
    • 负责人:
      Dawn Knight
    • 依托单位:
    Interactional variation online: harnessing emerging technologies in the digital humanities to analyse online discourse in different workplace contexts
    • 批准号:
      AH/W001608/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $40.17万
    • 财政年份:
      2021
    • 负责人:
      Dawn Knight
    • 依托单位: