EAGER: Automatic Speech Recognition for Uyghur
EAGER: Automatic Speech Recognition for Uyghur
批准号:
1519164
负责人:
Arienne Dwyer
金额:
$2.73万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-02-15 至 2016-08-31
中文摘要
语音工程的进步现在允许将音频转录为文本,甚至对于没有多少计算资源的语言也是如此。自动文本转录更多的语言允许公众,社区和研究人员访问以前无法访问的材料。该项目使用一种资源不足的语言进行数千小时的无线电广播,作为测试案例,以改进适用于任何语言的快速音频到文本开发技术。该项目允许语音工程师将技术应用于新语言,了解新语言的特点及其对语音识别性能的影响,以及如何克服这些特点,以建立更好的语音识别系统。它还使社区能够保存他们的语言,分发工具和数据,总的来说,改善他们的语言目前极度的资源限制。该项目鼓励学生在语音工程、语言学和新闻学等领域进行工作和思考。在这个EAGER项目中,维吾尔语(ISO 639-3: uig)是中亚新疆的一种资源严重不足的突厥语,大约有1100万使用者,用于测试自动语音识别(ASR)系统的快速发展,该系统的长期愿景是创建基于网络的语音和语言服务,包括发音字典生成、音频和文本数据存档以及词性标注。该项目是探索性的,因为该语言缺乏可计算的资源,但通过相关语言(土耳其语)进行引导有望快速发展ASR。该项目可以作为任何语言开发的模型,无论大小,并且具有潜在的变革性——首先,因为世界上许多语言都像维吾尔语一样,可用的计算资源很少。第二,因为许多纪实语言学家仍然完全依赖于非自动化的方法。
英文摘要
Advances in speech engineering now allow audio to be transcribed as text, even for languages for which there are few computational resources. Automating text transcription for more languages allows public, community, and researcher access to previously inaccessible materials. This project uses several thousand hours of radio broadcasts in an under-resourced language as a test case to improve rapid audio-to-text development techniques, which are applicable to any language. The project allows speech engineers to apply technology to new languages, to learn about the characteristics of new languages and their impact on speech recognition performance, and how to overcome them with the goal of building better speech recognition systems. It also enables communities to preserve their language, distribute tools and data, and overall, improve the current extreme resource limitations of their language. The project encourages students to work and think across the fields of speech engineering, linguistics and journalism.In this EAGER project, the Uyghur language (ISO 639-3: uig), a severely under-resourced Turkic language of Xinjiang in Central Asia with about 11 million speakers, is used to test the rapid development of an Automatic Speech Recognition (ASR) system with the long-term vision of creating web-based speech and language services including pronouncing dictionary generation, audio and text data archiving, and part-of-speech tagging. The project is exploratory because the language is devoid of computationally tractable resources, yet bootstrapping through a related language (Turkish) promises rapid ASR development. The project can serve as a model for such development for any language, large or small, and is potentially transformative -- first because so many of the world's languages are like Uyghur in having few available computational resources., and second because so many documentary linguists still rely entirely on non-automated methods.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CoLang: Institute for Collaborative Language Research
-
批准号:1065469
-
项目类别:Standard Grant
-
资助金额:$17.46万
-
财政年份:2011
-
负责人:Arienne Dwyer
-
依托单位:
Light Verbs in Uyghur
-
批准号:1053152
-
项目类别:Continuing Grant
-
资助金额:$32.8万
-
财政年份:2011
-
负责人:Arienne Dwyer
-
依托单位:
Interactive Inner Asia: documenting an endangered language contact area
-
批准号:1065524
-
项目类别:Standard Grant
-
资助金额:$25.92万
-
财政年份:2011
-
负责人:Arienne Dwyer
-
依托单位:
Collaborative Proposal: Workshop: Towards the Interoperability of Language Resources
-
批准号:0709732
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Arienne Dwyer
-
依托单位:
Conference:DT-Summit/Linguistics: Proposal for a Tool Development Workshop
-
批准号:0624048
-
项目类别:Standard Grant
-
资助金额:$2.5万
-
财政年份:2006
-
负责人:Arienne Dwyer
-
依托单位:
海外基金