课题基金 / 基金详情

Breaking the Unwritten Language Barrier

Breaking the Unwritten Language Barrier
打破不成文的语言障碍
批准号:
259117245
负责人:
Dr. Fatima Hamlaoui
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2014
资助国家:
德国
项目状态:
已结题
起止时间:
2013-12-31 至 2018-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
BLOB项目旨在借助自动语音和语言处理,特别是自动语音识别(ASR)和机器翻译(MT),支持非书面语言的记录。我们将讨论班图族三种主要是不成文的非洲语言(巴萨语、迈纳语和恩博西语)的记录。该项目的主要步骤是:1.以合理的成本收集语料库,采用三步法,遵循S·伯德和M·利伯曼的工作:收集社区中的大型语料库(100小时),包括引出的材料、故事、对话和广播;重新发言。由于录音的音质将是非常自然的,在嘈杂的环境中可能会有重叠的语音,由参考说话者仔细地重新发音将带来更准确的自动语音转录,并为语音/音系学研究提供更好的材料。翻译是记录一种新语言的自然方式;口头翻译将加速记录过程。我们的班图数据将被翻译成法语,法语是我们所研究社区的主要语言和第二语言。所收集的口头数据(班图语原文和法文译本)包含记录所研究语言的必要信息。预计ASR将自动生成源语言和目标语言的准确转录,并使用机器翻译在两者之间提供有意义的比对,以加快文件编制、描述和分析的主要任务。主要的自动处理步骤是:对所学习的语言进行音标。这一步骤首先需要一套与语言无关的电话模型,这些模型必须通过无人监督的适应技术来调整到所研究的语言;法语口语翻译的单词转录。需要调整语言和声学模型以获得高转录准确性;所研究语言的语音转录(原文、重写)之间的比对。对齐对于大规模声学-语音研究、语音和韵律数据挖掘以及方言变异研究是非常有价值的;跨语言对齐旨在将所研究语言中的音素序列与法语单词联系起来。这样的比对可能会被证明对词法研究、词汇和发音的阐述非常有用。该项目的成功有赖于语言学家和计算机科学家之间强大的德法合作。将通过一系列课程促进和加强合作,使科学界受益,而不是目前的联盟。在这些课程中,语言学家将向计算机科学家介绍记录一种未知语言的主要步骤,计算机科学家将介绍他们处理一种“新”语言的方法,从而生成语音转录和伪词对齐,并返回给语言学家。
英文摘要
The BULB project aims at supporting the documentation of unwritten languages with the help of automatic speech and language processing, in particular automatic speech recognition (ASR) and machine translation (MT). We will address the documentation of three mostly unwritten African languages of the Bantu family (Basaa, Myene and Embosi). The main steps of the project are:1. To collect the corpora at a reasonable cost, using a three step methodology, following the work of S. Bird and M. Liberman:collecting a large corpus of speech (100 hours) in a community, including elicited material, stories, dialogs and broadcasts;re-speaking. As the sound quality of the recordings will be very spontaneous, with possibly overlapping speech in noisy environments, carefully articulated re-speaking by a reference speaker will give rise to more accurate automatic phonetic transcriptions and to improved material for phonetic/phonological studies.oral translation. Translation is the natural way to document a new language; oral translations will accelerate the documentation process. Our Bantu data will be translated to French, a major language and a second language in the regions of our studied communities.2. The collected oral data (Bantu originals and French translations) contain the necessary information to document the studied languages. ASR is expected to automatically produce accurate transcriptions in source and target languages and MT to provide meaningful alignments between both, to speed up the major tasks of documentation, description and analysis. The major automatic processing steps are:phonetic transcription of the studied languages. This step requires first a set of language-independent phone models which must be tuned to the language under study via unsupervised adaptation techniques;word transcription of the oral French translations. Language and acoustic models need to be adapted to obtain high transcription accuracy;alignments between the phonetic transcriptions (originals, respeaking) of the studied language. Alignments are highly valuable for large scale acoustic-phonetic studies, phonological and prosodic data mining and dialectal variations studies;cross-language alignments that aim at linking phone sequences in the studied language with French words. Such alignments may prove very useful for morphological studies, vocabulary and pronunciation elaboration.The success of the project relies on a strong German-French cooperation between linguists and computer scientists. Cooperations will be fostered and strengthened by a series of courses benefiting the scientific community beyond the present consortium. During these courses, linguists will present to computer scientists the major steps to document an unknown language, and computer scientists will introduce their methods to process a "new" language thus generating phonetic transcriptions and pseudo-word alignments to be returned to linguists.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
Towards phoneme inventory discovery for documentation of unwritten languages
面向非书面语言记录的音素清单发现
DOI: 10.1109/icassp.2017.7953148
发表时间: 2017
期刊: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者: [Müller, Markus, Jörg Franke, Alex Waibel, Sebastian Stüker]
通讯作者: Sebastian Stüker
Neural Language Codes for Multilingual Acoustic Models
多语言声学模型的神经语言代码
DOI: 10.21437/interspeech.2018-1241
发表时间: 2018
期刊:
影响因子: --
作者: [Markus Müller, Sebastian Stüker, Alex Waibel]
通讯作者: Alex Waibel
Unsupervised Phoneme Segmentation of Previously Unseen Languages
以前未见过的语言的无监督音素分割
DOI: 10.21437/interspeech.2016-1440
发表时间: 2016
期刊:
影响因子: --
作者: [Vetter, Markus Müller, Fatima Hamlaoui, Graham Neubig, Satoshi Nakamura, Sebastian Stüker, Alex Waibel]
通讯作者: Alex Waibel
DBLSTM based multilingual articulatory feature extraction for language documentation
基于 DBLSTM 的语言文档多语言发音特征提取
DOI: 10.1109/asru.2017.8268966
发表时间: 2017
期刊: 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
影响因子: --
作者: [Müller, Markus, Sebastian Stüker, Alex Waibel]
通讯作者: Alex Waibel
共 6 条
    海外基金