Syntactic patterns in Pite Saami: A corpus-based exploration of 130 years of variation and change
Syntactic patterns in Pite Saami: A corpus-based exploration of 130 years of variation and change
批准号:
286335341
负责人:
Dr. Joshua Wilbur
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2016
资助国家:
德国
项目状态:
已结题
起止时间:
2015-12-31 至 2020-12-31
中文摘要
这个项目的最终目标是创建一个彻底的,基于语料库的描述的句法模式在皮特萨米语,一个高度濒危的乌拉尔语言在瑞典拉普兰。语料库将包括Pite Saami语的口语文本,代表从19世纪末到21世纪初一个多世纪的语言使用。考虑到这一点,源数据的年龄将被视为解释已证实模式变化的潜在因素,从而允许调查随时间的结构变化。具体而言,该项目将使用定量方法,试图回答以下主要研究问题。1.在Pite Saami语的各种短语和从句类型中,哪些成分是可能的和/或必需的?是否对某些结构有偏好?2.信息结构对构成结构有什么影响?3.语料库是否为句法模式的历时变化提供了证据?如果是,哪些模式受到影响?为了进行调查,必须首先创建注释语料库。为了有效地做到这一点,将改进现有的语言技术工具,以自动标记Pite Saami语的文本,包括词素、形态类别和词性。一本书长度的描述证明句法模式; 2.为一种濒临灭绝的萨米语建立一个完整注释的、跨越世纪的数字口语语料库,以供进一步研究; 3.使用语言技术工具自动注释濒危语言口语语料库的模型。计划中的句法描述将提供有关迄今为止未被描述的语言的新数据。这些数据不仅会引起乌拉尔语言学者的兴趣,特别是历史比较研究,而且会引起形式和功能方法的共时比较理论语言学家的兴趣。
英文摘要
The ultimate goal of this project is to create thorough, corpus-based descriptions of syntactic patterns in Pite Saami, a highly endangered Uralic language spoken in Swedish Lapland. The corpus will consist of Pite Saami texts in the spoken mode representing more than a century of language use from the late 19th up to the early 21st centuries. With this in mind, the age of source data will be treated as a potential factor in explaining variation in attested patterns, thus allowing for the investigation of structural changes through time.Specifically, the project will use quantitative methods in attempting to answer the following main research questions. 1. Which constituents are possible and/or required in the various Pite Saami phrase and clause types? Is there a preference for certain structures? 2. What effect does information structure have on constituent structure? 3. Does the corpus provide evidence for diachronic changes in syntactic patterns? If so, which patterns are affected?In order to carry out the investigation, an annotated corpus must first be created. To do this efficiently, extant language technology tools will be refined to automatically tag Pite Saami texts for lexeme, morphological categories and part-of-speech.The results of the project will be three-fold: 1. a book-length description of attested syntactic patterns; 2. a thoroughly annotated, digital spoken language corpus spanning more than a century of texts for an endangered Saami language, to be available for further research; and 3. a model for the use of language technology tools to automatically annotate a spoken language corpus for an endangered language.The planned syntactic description will provide new data concerning a hitherto under-described language. These data will not only be of interest to Uralic language scholars, particularly for historical-comparative studies, but also to synchronic comparative theoretical linguists with both formal and functional approaches.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
ELAN as a search engine for hierarchically structured, tagged corpora
ELAN 作为分层结构、标记语料库的搜索引擎
DOI:
10.18653/v1/w19-0308
发表时间:
2019
期刊:
Proceedings of the Fifth International Workshop on Computational Linguistics for Uralic Languages
影响因子:
--
作者:
[Wilbur]
通讯作者:
Wilbur
Envisioning Digital Methods for Fieldwork in the Arctic
设想北极实地考察的数字方法
DOI:
10.4324/9781003158295-22
发表时间:
2021
期刊:
影响因子:
--
作者:
[Partanen, Michael Rießler, J. Wilbur]
通讯作者:
J. Wilbur
DOI:
10.18653/v1/w17-0604
发表时间:
2017
期刊:
影响因子:
--
作者:
[C. Gerstenberger;N. Partanen;Michael Rießler;J. Wilbur]
通讯作者:
C. Gerstenberger;N. Partanen;Michael Rießler;J. Wilbur
Using Computational Approaches to Integrate Endangered Language Legacy Data into Documentation Corpora: Past Experiences and Challenges Ahead
使用计算方法将濒危语言遗产数据集成到文档语料库中:过去的经验和未来的挑战
DOI:
10.33011/computel.v2i.451
发表时间:
2019
期刊:
Proceedings of the Workshop on Computational Methods for Endangered Languages
影响因子:
--
作者:
[Blokland, Rogier, Niko Partanen, Michael Rießler, J. Wilbur]
通讯作者:
J. Wilbur
海外基金