Tibetan in Digital Communication: Corpus Linguistics and Lexicography
Tibetan in Digital Communication: Corpus Linguistics and Lexicography
批准号:
AH/J00152X/1
负责人:
Ulrich Pagel
金额:
$57.16万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2012
资助国家:
英国
项目状态:
已结题
起止时间:
2012 至 --
中文摘要
在时代、广度和流派的多样性上,西藏文学在各个方面都可以与英语相媲美。藏文字母是在公元650年发明的。目前可用的最早的可安全标注日期的文件可以追溯到公元763年左右。从那时起,文学创作有增无减,一直延续到今天。然而,藏语的词典编纂资源非常不足,远远低于说英语的人所能获得的资源。学习藏语的学生总共可以查阅十几本词典,其中大部分是古藏语词典。这些词汇的范围往往定义不清,而且没有一个符合科学词典编纂的标准。此外,没有一部作品涵盖了藏族文学最早的时期--《旧西藏》(公元650-1000年)。我们建议建立的语料库和工具将作为推进编写一本类似牛津英语词典的全面历史藏语词典的第一步。为了实现这一目标,我们建议建立一个涵盖藏语整个历史的大型藏语语料库,摘自古藏语、古藏语和现代藏语。过去,为了编纂大型词典,学者们使用整理和存储在巨大档案柜中的繁重的纸简集合。计算语言学的进步意味着,现在可以通过创建带注释的数字语料库来更彻底、更有效地完成这项工作。但是,一旦我们的语料库经过仔细的分析和标记,不仅将为编纂迄今难以想象的藏语词典铺平道路,而且还将为其他广泛的重要研究项目奠定基础。通过将其安装在Web上,来自广泛学科(历史、宗教、文学、语言学等)的学者使用藏语材料将能够搜索它,并将其内容用于自己的研究。因此,它很可能成为一系列研究计划的基础,使学术界的许多不同群体受益。在学术界之外,在现代电子通信世界中,我们的语料库将为创造新的藏语数字技术(短信、自动翻译等)奠定基础。开发语言软件所需的高额投资使没有商业或政治权力的语言孤立,资源匮乏。数字通信技术是建立在基本的语言处理工具(例如,分词程序、词性标记)的基础上的,而这正是我们打算创造的类型。我们的工作将降低开发此类技术的成本,从而吸引商业兴趣。尽管有200多万人说藏语,但在电子媒体上,藏语几乎没有作为一种口语出现。我们试图通过创建一种电子资源来补救这一点,该资源将恢复藏人的选择,无论他们的居住地或被接纳的国籍如何,在一个日益由数字通信塑造的世界中,他们可以选择使用他们的语言。
英文摘要
In age, breadth and diversity of genre, Tibetan literature is in every way comparable to English. The Tibetan alphabet was invented in 650 CE. The earliest currently available securely dateable document dates to ca. 763 CE. Literary production has continued from that time unabated until today. Yet, the lexicographical resources of Tibetan are very inadequate and vastly inferior to what is available to English speakers. In total, students of Tibetan can draw on about a dozen dictionaries, most for Classical Tibetan. The scope of these lexicons tends to be poorly defined, and none of them meets the standards of scientific lexicography. Moreover, there is not a single work that covers the earliest period of Tibetan literature, Old Tibetan (650-1000 CE). The corpus and tools we propose to create will serve as the first step to advance the compilation of a comprehensive historical Tibetan dictionary akin to the Oxford English Dictionary.In order to achieve this, we propose to produce a large corpus of Tibetan texts spanning the language's entire history, drawn from Old, Classical and Modern Tibetan. In the past, scholars used laborious collections of slips organised and stored in vast filing cabinets in order to compile large dictionaries. Advances in computational linguistics mean that this work can now be achieved more thoroughly and effectively through the creation of annotated digital corpora. But our corpus, once carefully analysed and tagged, will not only pave the way for the compilation of Tibetan dictionaries of hitherto inconceivable calibre, but it will also prepare the ground for a wide range of other significant research initiatives. By mounting it on the Web, scholars from a wide range of disciplines (history, religion, literature, linguistics, etc.) working with Tibetan language materials will be able to search it and use its content for their own research. It is thus likely to become foundational to a vast array of research initiatives, benefiting many different constituencies in academia.Outside academia, in the modern world of electronic communication, our corpus will lay the foundation for the creation of new digital technologies for Tibetan (text messaging, automated translation, etc.). The high investment required to develop language software leaves languages without commercial or political power isolated and poorly resourced. Digital communication technologies are built on basic language processing tools (eg, word-segmentation programmes, part-of-speech taggers) of the very type we propose to create. Our work will reduce the cost to develop such technologies and thus attract commercial interest. Although Tibetan is spoken by more than two million people, it is barely represented in electronic media as a spoken language. We seek to remedy this by creating an electronic resource that will restore to Tibetans, irrespective of their residence or adopted nationality, the choice to use their language as they see fit in a world that is increasingly shaped by digital communication.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2015
期刊:
Rocznik Orientalistyczny
影响因子:
--
作者:
[Hill, N]
通讯作者:
Hill, N
Tibetan vlan 'reply'
藏语 vlan 回复
DOI:
10.1017/s1356186314000455
发表时间:
2014
期刊:
Journal of the Royal Asiatic Society of Great Britain & Ireland
影响因子:
--
作者:
[HILL N]
通讯作者:
HILL N
Disambiguating Tibetan verb stems with matrix verbs in the indirect infinitive construction
间接不定式结构中用矩阵动词消除藏语动词词干的歧义
DOI:
--
发表时间:
2015
期刊:
Bulletin of Tibetology
影响因子:
--
作者:
[Garrett, E]
通讯作者:
Garrett, E
Constituent Order in the Tibetan Noun Phrase
藏语名词短语的构成顺序
DOI:
--
发表时间:
2015
期刊:
SOAS Working Papers in Linguistics
影响因子:
--
作者:
[Garrett, E]
通讯作者:
Garrett, E
A Rule-based Part-of-speech Tagger for Classical Tibetan
基于规则的古典藏语词性标注器
DOI:
10.5070/h913224023
发表时间:
2014
期刊:
Himalayan Linguistics
影响因子:
--
作者:
[Garrett E]
通讯作者:
Garrett E
共 8 条
Lexicography in Motion: A History of the Tibetan Verb
-
批准号:AH/P004644/1
-
项目类别:Research Grant
-
资助金额:$100.99万
-
财政年份:2017
-
负责人:Ulrich Pagel
-
依托单位:
Locating Culture, Religion and the Self: A Study of the Tantric Community in Rebkong (East Tibet)
-
批准号:AH/F009216/1
-
项目类别:Research Grant
-
资助金额:$33.81万
-
财政年份:2008
-
负责人:Ulrich Pagel
-
依托单位:
国内基金
海外基金
登录
查看更多内容
超灵敏高分辨的Digital-CRISPR技术用于免扩增的多重核酸检测
-
批准号:22104048
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:陈勇
-
依托单位:
基于Digital Twin的数控机床智能运行维护方法研究
-
批准号:51875323
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2018
-
负责人:胡天亮
-
依托单位:
基于数字PCR(digital-PCR)技术的耳聋无创产前检测研究
-
批准号:LQ19H040016
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2018
-
负责人:严恺
-
依托单位:
基于Digital LAMP技术的循环肿瘤细胞检测和分型新方法研究
-
批准号:81702102
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:王纪东
-
依托单位:
基于表面工程的外泌体digital PCR定量分析体系的构建及转化医学研究
-
批准号:81702959
-
项目类别:青年科学基金项目
-
资助金额:10.0万元
-
批准年份:2017
-
负责人:田庆常
-
依托单位: