Integrating, Disseminating, and Archiving Components of the Shoshoni Language Project
Integrating, Disseminating, and Archiving Components of the Shoshoni Language Project
批准号:
1911603
负责人:
Marianna Di Paolo
金额:
$19.74万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
未结题
起止时间:
2019-08-15 至 2025-01-31
中文摘要
美国国会1990年通过的《美洲原住民语言法案》承认了美国原住民语言的独特地位和价值。Shoshoni[ISO 639-3 shh]是从怀俄明州到中美洲所说的乌托-阿兹特克语系的最北端的成员。今天,肖肖尼语仍然是戈舒特和肖肖尼部落特征的重要组成部分。在1960年‘S-1970年’S,已故的犹他大学的威克·R·米勒从几个不同的品种录制了说肖肖尼语(出生于1875年至1920年)的录音,代表了大盆地语言中最广泛的纪录库,对西部各州的几个部落社区具有至关重要的文化、历史和语言重要性。过去对Shoshoni的语言学研究主要集中在孤立的句子内部结构和词的结构上,而本项目将重点放在它的语音系统和语篇层次结构上。更广泛的影响包括从万豪图书馆(犹他大学)和加州语言档案馆(加州大学伯克利分校)以免费在线资源的形式提供这两个语料库。该项目还将为来自Shoshoni部落社区的本科生提供关于计算语言研究项目的宝贵经验,并加强这些年轻人与在该项目上合作的两名讲母语的长老之间的互动。该团队还将制作维克·R·米勒收藏集中传统故事子集的印刷版和易于阅读的电子版,并将它们传播给三个合作参与该项目的社区:特-莫克部落的南福克乐队委员会、戈舒特保留地的邦联部落和伊利·肖肖尼部落。虽然肖肖尼语作为一种美洲原住民语言有相当好的文献记载,但它的语篇结构以及语音和音系学相对研究较少。因此,这些重大差距将通过开发两个语料库来弥补。首先,将对36个故事进行标记,以产生一个对句子级和语篇级语言学研究有价值的电子可搜索数据库。第二,一个语音和语音有价值的语料库,包括音频-文本网格对单词和句子大小的录音,这些录音将强制对齐并进行微调。在生成的语料库中,代表每个元音和辅音的音素将与声音文件的相应部分对齐,使研究人员能够自动对每个声音进行声学语音分析。这种文本到音频对齐的语料库已经存在于大多数语言,如英语、德语、日语和西班牙语,这使得它们的声音系统相对容易学习,从而导致了能够快速处理口语的电子产品的开发。这些大多数语言语料库是使用被称为强制对齐的昂贵的、特定于语言的计算工具来准备的。我们的项目将训练蒙特利尔强制校准器将4,000-5,000个Shoshoni单词和短句的文本进行发音。这样做将提供一个模型,说明如何以较低的成本使用通用的强制对齐器来对齐任何小型、未被研究的语言的文本到音频数据。由此产生的强制对齐的Shoshoni语料库将极大地加快对这种语音复杂语言的声学分析,并导致许多相对便宜但深入、科学合理的研究。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The Native American Languages Act, passed by the U.S. Congress in 1990, recognizes the unique status and value of Native American languages. Shoshoni [ISO 639-3 shh] is the northernmost member of the Uto-Aztecan language family, languages spoken from Wyoming to Central America. The Shoshoni language today continues to be an important component of Goshute and Shoshone tribal identity. In the 1960's-1970's, the late Wick R. Miller, of the University of Utah, taperecorded speakers of Shoshoni (born from ~1875-1920) from several different varieties, representing the most extensive documentary corpus of any Great Basin language, of vital cultural, historical, and linguistic importance to several tribal communities in the Western states. Past linguistic studies of Shoshoni have largely focused on the internal structure of sentences in isolation and on the structure of words, while this project will focus on its sound system and discourse-level structure. Broader impacts include the availability of the two corpora as free online resources from the Marriott Library (University of Utah) and the California Language Archive (UC-Berkeley). The project will also provide undergraduates from Shoshoni-speaking tribal communities with valuable experience on a computational linguistic research project, and enhance interactions between these young people and the two native-speaker elders collaborating on the project. The team will also produce a print version and an easy-to-read electronic version of a subset of the traditional stories from the Wick R. Miller Collection and disseminate them to the three communities collaborating on the project, the South Fork Band Council of the Te-Moak Tribe, the Confederated Tribes of the Goshute Reservation and the Ely Shoshone Tribe.While Shoshoni is fairly well-documented for a Native American language, its discourse structure and its phonetics and phonology are relatively understudied. Thus, these significant gaps will be remedied by the development of two corpora. First, the 36 stories will be marked up to produce a electronically-searchable database valuable for sentence-level as well as discourse-level linguistic studies. Second, a phonological and phonetically valuable corpus, consisting of audio-TextGrid pairs of word and sentence-sized recordings which will be force aligned and fine-tuned. In the resulting corpus, the phonemes representing each vowel and consonant will be aligned with the corresponding part of the sound file, allowing researchers to automate the acoustic phonetic analysis of each sound. Such text-to-audio aligned corpora already exist for majority languages such as English, German, Japanese, and Spanish, making their sound systems relatively easy to study and thus leading to the development of electronic products that can quickly process spoken language. These majority language corpora are prepared using costly, language-specific computational tools called forced aligners. Our project will train the Montreal Forced Aligner to align the text of 4,000-5,000 Shoshoni words and short sentences to sound. Doing so will provide a model of how to inexpensively use a generic forced aligner to align text-to-audio data for any small, understudied language. The resulting forced-aligned Shoshoni corpus will greatly speed up the acoustic analysis of this phonologically complex language and lead to many relatively inexpensive, but in-depth, scientifically-sound research studies.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Doctoral Dissertation Research: Linguistic Variation as a Marker of Ethnic Identity in High School Setting
-
批准号:1749582
-
项目类别:Standard Grant
-
资助金额:$0.93万
-
财政年份:2018
-
负责人:Marianna Di Paolo
-
依托单位:
Workshop on Sociophonetic Methodology for the 2011 Linguistics Institute
-
批准号:1058778
-
项目类别:Standard Grant
-
资助金额:$4.25万
-
财政年份:2011
-
负责人:Marianna Di Paolo
-
依托单位:
海外基金