课题基金 / 基金详情

Lexicography in Motion: A History of the Tibetan Verb

Lexicography in Motion: A History of the Tibetan Verb
动态词典编纂:藏语动词史
批准号:
AH/P004644/1
负责人:
Ulrich Pagel
金额:
$100.99万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

Ulrich Pagel的其他基金

相似基金

相关文献

中文摘要
翻译
在某种程度上,大多数语言使用者都依赖词典作为词典编纂信息的权威来源。第一个有记录的字典可以追溯到苏美尔时代(公元前3千年),在阿卡德语和苏美尔语之间的语言融合过程中(“Sprachbund”)编纂。从那时起,字典在跨文化交流和跨语言、跨民族国家的先进科学研究中发挥了关键作用。现代词典编纂仍然服务于这些目标,但其方法已经变得面目全非。卡片目录已经让位于数据库和数字资源,它们提供了访问更大的语言数据池的途径。今天,几乎所有的词典编纂者都部署了文本语料库和语料库查询工具,以提高定义的经验基础,并为最终用户提供上下文示例。我们建议利用这些发展来创造一个基于语料库的藏语动词历时词汇。动词在大多数句子中起着中心作用。对一个动词的意义的了解导致了它所需要的论证,以及这些论证所扮演的语义角色。我们的词典借鉴了这些联系。它将允许用户推断出一个句子的完整结构,主要基于结束动词和伴随参数的类型。此外,本文还描绘了从公元8世纪最早的藏文记录到当代的动词的形态和语义变化。每个动词在语料库中追溯到其早期在古藏文材料中的出现,然后与其在古典藏文和现代藏文中的应用进行比较。现有的一些词典包含零星的历时信息,但这些信息从未被分析或与其他数据并列。我们建议以系统的方式对历时证据进行识别、检查和背景化,以便更好地掌握藏语的整体演变。语料库资源和处理工具是现代词典编纂中不可缺少的组成部分。对于藏语来说,其中一些工具现在是可用的。“数字传播中的藏语”项目收集了大量藏语材料,并使用词性标注,涵盖古藏语、古典藏语和现代藏语。对于词典,我们通过在自然语言处理(NLP)软件上运行一系列自动查询来挖掘其内容。首先,我们创建一个内部工作流工具。这使我们能够系统而全面地对语料库中的所有藏语动词进行分类。然后,动词的不同形式被分组在不同的条目中;我们深入分析了语义歧义、重复变化或形态不规则的词干。同时,我们识别并标记与每个动词相连的实参。我们单独或累积地使用这些数据来生成词典的引用和定义。动词词汇将成为学生和学者们研究西藏文化的任何一个方面,无论是过去还是现在,不可或缺的财富。在学术界之外,通过其现代组成部分,词典改善了在政治和经济领域获得与西藏有关内容的途径。发展援助、人道主义援助、医疗服务和教育支助最好在与受援国对话的情况下提供。这些对话必须用藏语进行。我们开发的软件也为藏语使用者提供了新的数字工具。IT行业不愿投资于一个几乎没有政治或经济影响力的民族的语言。它的使用者被排除在网络的巨大资源之外。这些技术的关键是基本语言资源工具包(BLARK)的可用性。词汇和谓词软件使我们离完成藏文BLARK又近了一步。
英文摘要
At one point or another, most language users rely on dictionaries as authoritative source of lexicographical information. The first recorded dictionaries date back to Sumerian times (3rd millennium BCE) compiled in the course of the linguistic convergence ('Sprachbund') between Akkadian and Sumerian. Since then, dictionaries have played a key role in intercultural communication and advanced scientific research across languages and nation states. Modern-day lexicography still serves these goals, but its methods have changed beyond recognition. Card catalogues have given way to databases and digital resources that offer access to a much larger pool of linguistic data. Today, practically all lexicographers deploy text corpora and corpus querying tools, both to sharpen the empirical base of definitions and to provide contextual examples for the end user. We propose to take advantage of these developments to create a corpus-based diachronic lexicon of Tibetan verbs.Verbs play a central role in most sentences. Knowledge of the meaning of a verb leads to the arguments it requires and to the semantic roles the arguments, in turn, assume. Our lexicon draws on these links. It will allow the user to infer the complete structure of a sentence, based primarily on the terminal verb and the type of accompanying arguments. Furthermore, it charts the morphological and semantic changes of the verbs from the earliest records of Tibetan in the 8th century CE to contemporary times. Each verb is tracked to its earlier occurrence in the Old Tibetan material within the corpus and then compared with its applications in Classical and Modern Tibetan. Some of the existing dictionaries contain sporadic diachronic information, but this is never analysed or juxtaposed with other data. We propose to identify, examine and contextualise the diachronic evidence in a systematic fashion in order to obtain a better grasp of the evolution of the Tibetan language overall.Corpus resources and processing tools constitute indispensable components in modern lexicography. For Tibetan, some of these tools are now available. 'Tibetan in Digital Communication' produced a large corpus of Tibetan language material, with part-of-speech-tagging, spanning Old, Classical and Modern Tibetan. For the lexicon, we mine its content by running a series of automated queries drawing on Natural Language Processing (NLP) software. At first, we create an internal workflow tool. This allows us to categorise, both systematically and comprehensively, all the Tibetan verbs within the corpus. The different forms of the verbs are then grouped together in discrete entries; we analyse in depth verbal stems that display semantic ambiguity, repeated change or morphological irregularity. In parallel, we identify and label the arguments connected with each verb. We use this data, individually and cumulatively, to generate the citations and definitions for the lexicon.The verb lexicon will become an indispensable asset for students and scholars alike, working on any one of the many facets of Tibetan culture, past and present. Outside academia, through its modern component, the lexicon improves access to Tibet-related content in the political and economic sphere. Development aid, humanitarian assistance, medical provisions and educational support are best delivered in conversation with the recipients. These conversations must be conducted in Tibetan. Very few Tibetans are fluent in English and most do not wish to communicate in Chinese. The software we create also advances the creation of new digital tools for Tibetan speakers. The IT sector is reluctant to invest in the language of a people that holds little political or economical influence. Its speakers are excluded from the vast resources of the web. Key to such technologies is the availability of a Basic Language Resource Kit (BLARK). The lexicon and predicate software bring us one step closer to the completion of a BLARK for Tibetan.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
NER for Tibetan and Mongolian Newspapers
藏文和蒙古文报纸的 NER
DOI: 10.33774/coe-2021-xhw9l
发表时间: 2021
期刊:
影响因子: --
作者: [Barnett R]
通讯作者: Barnett R
A CG3 Constraint Grammar to detect verb dependencies for the Classical Tibetan Language
用于检测古典藏语动词依赖性的 CG3 约束语法
DOI: --
发表时间:
期刊: Himalayan Linguistics
影响因子: --
作者: [Faggionato, C]
通讯作者: Faggionato, C
Dictionaries as collections of data stories: an alternative post-editing model for historical corpus lexicography
字典作为数据故事的集合:历史语料库词典编纂的另一种译后编辑模型
DOI: --
发表时间: 2021
期刊: eLex
影响因子: --
作者: [Ligeia Lugli]
通讯作者: Ligeia Lugli
Lexicography in Motion
运动中的词典编纂
DOI: --
发表时间: 2019
期刊: Jahrbucheintrag BAdW
影响因子: --
作者: [Rode, Samyo]
通讯作者: Rode, Samyo
共 10 条
    Tibetan in Digital Communication: Corpus Linguistics and Lexicography
    Locating Culture, Religion and the Self: A Study of the Tantric Community in Rebkong (East Tibet)
    海外基金