Cross-linguistic phonetics and morphology using a time-aligned multilingual reference corpus built from documentations of 50 languages: Big data on small languages
Cross-linguistic phonetics and morphology using a time-aligned multilingual reference corpus built from documentations of 50 languages: Big data on small languages
批准号:
411066783
负责人:
Privatdozent Dr. Frank Seifart, since 11/2019
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2019
资助国家:
德国
项目状态:
已结题
起止时间:
2018-12-31 至 2022-12-31
中文摘要
语速和停顿为我们了解人类语言产生系统的认知-神经和生理-发音基础提供了一个窗口,但这一领域的跨语言差异仍未得到充分研究。这个项目通过对50种不同语言的自发口语样本进行比较研究,填补了这一空白。为此,我们创建了语言文档数据(DoReCo)的多语言参考语料库,由注释和相关的音频记录组成,这些注释和音频记录被归档在诸如The language Archive (TLA)之类的存储库中,特别是来自DOBES集合。DoReCo将从已经转录的数据中构建,翻译成主要语言,并在话语单位级别与音频文件进行时间对齐。在当前的项目中,这些数据将在音素级别按时间排列。我们已经确定了至少50种语言,其中至少有10,000个单词的语料库可以包含在DoReCo中,并且至少有30种语言的子集已经对语素中断和语素注释。在DoReCo中,子语料库和注释被视为可引用的出版物,提供永久标识符并与CC BY 4.0许可相关联。DoReCo作为一个平台,可以方便地访问来自50多种语言的100多万字的带注释的语料库数据,用于口语的跨语言研究,它将产生超出DoReCo项目特定研究目标的持久影响。这是对全球语言多样性和文化遗产的开放、可复制科学的前所未有的贡献。DoReCo的两个具体研究目标都解决了人类语言的普遍性限制,这些限制来自于物种范围内的发音和认知特性:首先,我们研究语音延长的模式,目的是在以下方面建立通用模式与语言特定模式:(i)不同类型的语音片段在持续时间上的变化程度(例如元音与不同类型的辅音)——反映发音和感知限制;(ii)词尾延长作为主要韵律边界与次要韵律边界的指示——反映规划和潜在信号话语单位的认知限制。其次,我们研究了语素在时间分布上的普遍模式和语言特定模式,包括(i)每秒语素的信息率和(ii)停顿间单位的语素数量,这两者都反映了语言使用的认知限制。该项目将由一个跨学科团队进行,汇集了文献语言学、语音学、类型学和数量语言学方面的专业知识,并得到德国和法国两个领先研究中心的大力支持。
英文摘要
Speech rate and pauses provide us with a window into the cognitive-neural and physiological-articulatory bases of the human language production system, but crosslinguistic variation in this domain remain understudied. This project fills this gap by comparative studies of spontaneously spoken language in a diverse sample of 50 languages. For this purpose, we create a multilingual reference corpus of language documentation data (DoReCo) consisting of annotations and associated audio recordings that are archived at repositories such as The Language Archive (TLA), especially from the DOBES collection. DoReCo will be built from data that are already transcribed, translated into a major language, and time-aligned at the level of discourse units with audio files. Within the current project, these data will be time-aligned at the phoneme level. We have identified at least 50 languages, from which corpora of at least 10,000 words can be included in DoReCo, and a subset of at least 30 of these, which are additionally already annotated for morpheme breaks and morpheme glosses. In DoReCo, subcorpora and annotations are treated as citable publications, provided with a permanent identifier and associated with a CC BY 4.0 license. DoReCo will have a lasting effect beyond the specific research goals of the DoReCo project, as a platform for easy access to over one million words of annotated corpus data from over 50 languages for cross-linguistic research on spoken language. This represents an unprecedented contribution to open, reproducible science regarding global linguistic diversity and cultural heritage. Both of DoReCo’s two specific research goals address the universality of constraints on human language arising from species-wide articulatory and cognitive properties: Firstly, we investigate patterns of phonetic lengthening with the aim towards establishing universal vs. language-specific patterns in (i) the degree to which different types of phonological segments undergo variation in duration (e.g. vowels vs. different types of consonants)–reflecting articulatory and perceptual constraints–and (ii) word-final lengthening as indicative of major vs. minor prosodic boundaries–reflecting cognitive constraints on planning and potentially signalling discourse units. Secondly, we investigate universal vs. language-specific patterns in the temporal distribution of morphemes regarding (i) information rate in terms of morphemes per second and (ii) the number of morphemes in inter-pausal units–both reflecting cognitive constraints on language use. The project will be carried out by an interdisciplinary team bringing together expertise on documentary linguistics, phonetics, typology, and quantitative linguistics, with strong institutional support from two leading research centres in Germany and France.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金