课题基金 / 基金详情

Building a corpus of phonemic lexicons to study information theoretic universals

Building a corpus of phonemic lexicons to study information theoretic universals
建立音素词典语料库来研究信息论共性
批准号:
1829290
负责人:
Uriel Cohen Priva
金额:
$39.11万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2024-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Human language use reflects the nature of human communication. For instance, frequent words tend to have fewer sounds than infrequent ones, which facilitates quick production and understanding. However, little is known about more fine-grained distinctions. For instance, English has more /k/ than /p/ sounds. Does that reflect a property of human language and its physiological and perceptual nature or a historical accident? Answering such questions requires comparative data on the frequency and phonological makeup of words in many languages. This project will build on existing textual sources and word frequency lists to provide the phonological makeup of words in close to 200 low-resource languages. The phonological word lists will provide an invaluable resource to the understanding of human language and provide much-needed linguistic resources to low-resource languages. The outputs of the project will be made public and easily accessible, thereby assisting in documenting and teaching the processed languages, and in building computational linguistic resources such as text-to-speech engines. The research team, including trained undergraduate and graduate students, will create rules to translate alphabets to phonemic representation for multiple languages. The team will then collect textual resources and word frequency lists from publicly available sources such as online Bibles, newspapers, and movie subtitles. The rules will be applied separately to each source and the resulting phonological representations will be made publicly available, such that not only researchers but also the general public will be able to use and interact with the data. The researchers will proceed to use the data to investigate whether the information theoretic properties of sounds have distributional universality: do sounds tend to provide similar amounts of information cross-linguistically, and if so, does their information content correlate with their phonetic properties? Universality is an age-old question, and the similarities and differences of properties across language can provide new insights into language use. Specifically, the researchers will use information theoretic properties to predict whether low information or other previously studied phonological properties are likely to promote consonant weakening in those languages.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Schwa’s duration and acoustic position in American English
美式英语中 Schwa 的持续时间和声学位置
DOI: 10.1016/j.wocn.2022.101198
发表时间: 2023
期刊: Journal of Phonetics
影响因子: 1.9
作者: [Cohen Priva, Uriel, Strand, Emily]
通讯作者: Strand, Emily
The stability of segmental properties across genre and corpus types in low-resource languages
低资源语言中跨流派和语料库类型的分段属性的稳定性
DOI: 10.7275/fttf-fq95
发表时间: 2020
期刊: Proceedings of the Society for Computation in Linguistics
影响因子: --
作者: [Cohen Priva, Uriel, Yang, Shiying, Strand, Emily]
通讯作者: Strand, Emily
海外基金