Lexicography in Motion: A History of the Tibetan Verb
Lexicography in Motion: A History of the Tibetan Verb
批准号:
AH/P004644/1
负责人:
Ulrich Pagel
金额:
$100.99万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
At one point or another, most language users rely on dictionaries as authoritative source of lexicographical information. The first recorded dictionaries date back to Sumerian times (3rd millennium BCE) compiled in the course of the linguistic convergence ('Sprachbund') between Akkadian and Sumerian. Since then, dictionaries have played a key role in intercultural communication and advanced scientific research across languages and nation states. Modern-day lexicography still serves these goals, but its methods have changed beyond recognition. Card catalogues have given way to databases and digital resources that offer access to a much larger pool of linguistic data. Today, practically all lexicographers deploy text corpora and corpus querying tools, both to sharpen the empirical base of definitions and to provide contextual examples for the end user. We propose to take advantage of these developments to create a corpus-based diachronic lexicon of Tibetan verbs.Verbs play a central role in most sentences. Knowledge of the meaning of a verb leads to the arguments it requires and to the semantic roles the arguments, in turn, assume. Our lexicon draws on these links. It will allow the user to infer the complete structure of a sentence, based primarily on the terminal verb and the type of accompanying arguments. Furthermore, it charts the morphological and semantic changes of the verbs from the earliest records of Tibetan in the 8th century CE to contemporary times. Each verb is tracked to its earlier occurrence in the Old Tibetan material within the corpus and then compared with its applications in Classical and Modern Tibetan. Some of the existing dictionaries contain sporadic diachronic information, but this is never analysed or juxtaposed with other data. We propose to identify, examine and contextualise the diachronic evidence in a systematic fashion in order to obtain a better grasp of the evolution of the Tibetan language overall.Corpus resources and processing tools constitute indispensable components in modern lexicography. For Tibetan, some of these tools are now available. 'Tibetan in Digital Communication' produced a large corpus of Tibetan language material, with part-of-speech-tagging, spanning Old, Classical and Modern Tibetan. For the lexicon, we mine its content by running a series of automated queries drawing on Natural Language Processing (NLP) software. At first, we create an internal workflow tool. This allows us to categorise, both systematically and comprehensively, all the Tibetan verbs within the corpus. The different forms of the verbs are then grouped together in discrete entries; we analyse in depth verbal stems that display semantic ambiguity, repeated change or morphological irregularity. In parallel, we identify and label the arguments connected with each verb. We use this data, individually and cumulatively, to generate the citations and definitions for the lexicon.The verb lexicon will become an indispensable asset for students and scholars alike, working on any one of the many facets of Tibetan culture, past and present. Outside academia, through its modern component, the lexicon improves access to Tibet-related content in the political and economic sphere. Development aid, humanitarian assistance, medical provisions and educational support are best delivered in conversation with the recipients. These conversations must be conducted in Tibetan. Very few Tibetans are fluent in English and most do not wish to communicate in Chinese. The software we create also advances the creation of new digital tools for Tibetan speakers. The IT sector is reluctant to invest in the language of a people that holds little political or economical influence. Its speakers are excluded from the vast resources of the web. Key to such technologies is the availability of a Basic Language Resource Kit (BLARK). The lexicon and predicate software bring us one step closer to the completion of a BLARK for Tibetan.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
NER for Tibetan and Mongolian Newspapers
藏文和蒙古文报纸的 NER
DOI:
10.33774/coe-2021-xhw9l
发表时间:
2021
期刊:
影响因子:
--
作者:
[Barnett R]
通讯作者:
Barnett R
A CG3 Constraint Grammar to detect verb dependencies for the Classical Tibetan Language
用于检测古典藏语动词依赖性的 CG3 约束语法
DOI:
--
发表时间:
期刊:
Himalayan Linguistics
影响因子:
--
作者:
[Faggionato, C]
通讯作者:
Faggionato, C
Dictionaries as collections of data stories: an alternative post-editing model for historical corpus lexicography
字典作为数据故事的集合:历史语料库词典编纂的另一种译后编辑模型
DOI:
--
发表时间:
2021
期刊:
eLex
影响因子:
--
作者:
[Ligeia Lugli]
通讯作者:
Ligeia Lugli
Lexicography in Motion
运动中的词典编纂
DOI:
--
发表时间:
2019
期刊:
Jahrbucheintrag BAdW
影响因子:
--
作者:
[Rode, Samyo]
通讯作者:
Rode, Samyo
Smart lexicography for under-resourced languages
资源贫乏语言的智能词典编目
DOI:
--
发表时间:
2019
期刊:
影响因子:
--
作者:
[Ligeia Lugli]
通讯作者:
Ligeia Lugli
共 10 条
Tibetan in Digital Communication: Corpus Linguistics and Lexicography
-
批准号:AH/J00152X/1
-
项目类别:Research Grant
-
资助金额:$57.16万
-
财政年份:2012
-
负责人:Ulrich Pagel
-
依托单位:
Locating Culture, Religion and the Self: A Study of the Tantric Community in Rebkong (East Tibet)
-
批准号:AH/F009216/1
-
项目类别:Research Grant
-
资助金额:$33.81万
-
财政年份:2008
-
负责人:Ulrich Pagel
-
依托单位:
海外基金