课题基金 / 基金详情

基于深度迁移学习的安多藏语语音识别研究

批准号:
62066039
项目类别:
地区科学基金项目
资助金额:
36.0 万元
负责人:
黄鹤鸣
依托单位:
学科分类:
模式识别与数据挖掘
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
黄鹤鸣

项目摘要

结项摘要

黄鹤鸣的其他基金

相似基金

相关文献

中文摘要
安多藏语是藏语的三大主体方言之一,从语音识别角度看,属于低资源或者极低资源语言。关于安多藏语语音识别,仅有个别学者进行了探索性研究,研究内容分散且不够深入。因此,项目组拟开展基于深度迁移学习的安多藏语语音识别研究。本项目主要进行以下四个方面的研究:第一,进一步完善已有的安多藏语语音样本数据库;第二,从藏语为低资源语言的现状出发,研究基于深度迁移学习的安多藏语语音特征提取;第三,构建基于注意机制、对抗机制的深度神经网络声学模型;第四,构建基于Transformer的语言模型。本项目主要利用深度迁移学习、注意机制、对抗机制等最新模式识别技术,重点解决低资源条件下面向安多藏语的特征提取、声学模型构建、语言模型构建等关键科学问题,构建高效、准确的安多藏语语音识别系统。积极开展本项目研究有利于丰富低资源语言语音识别研究,实现藏语的语音识别输入,完善藏语的人机交互,保证网络中传播的藏语信息内容的安全。
英文摘要
Tibetan is a low resource language and Amdo Tibetan is one of its three main dialects. Up to the present, only a few researchers have done some exploratory researches on Amdo dialect speech recognition, and there is no systematical, effective, and practical Tibetan speech recognition system. Therefore, this research team will devote itself to this challenging project. Firstly, the team will construct a speech database of Tibetan Amdo dialect. Secondly, the team will explore some deep transfer learning-based feature extraction methods for low-resource language such as Amdo Tibetan. Thirdly, the team will construct an effective acoustic model based on attention mechanism and adversarial mechanism. Fourthly, the team will construct a Transformer-based language model for Tibetan..This research team will mainly focus on feature extraction, acoustic modelling, and language modelling for low resource language such as Amdo Tibetan by enrolling such novel techniques as deep transfer learning, attention mechanism, adversarial mechanism, etc. In this way, the team could set up a systematic, effective, and practical speech recognition system for Amdo Tibetan..The study of this project have many great theoretical and social significances. It can benefit the theory development of low resource language speech recognition, can perfect human-machine interface related to Tibetan language, and can ensure the security of Tibetan language information in internet. Furthermore, this project can train a group of talented researchers for Tibetan speech recognition.
安多语是藏语的三大主体方言之一,从语音识别角度看,属于低资源或者极低资源语言。关于安多藏语语音识别,仅有个别学者进行了探索性研究,研究内容分散且不够深入。受本项目资助,项目组主要开展了以下四个方面的研究:第一,构建了安多藏语语音样本数据库,为藏语语音识别研究奠定基础;第二,利用端到端的语音识别技术,将声学模型、语言模型和解码单元等模块集成到一个单一的神经网络中,解决传统语音识别模型中各模块独立、模型不能联合优化的问题,大大降低了语音识别模型的复杂性;第三,构建了一个基于LAS模型的安多藏语语音识别系统;第四,研究了并提出了一些有效的语音增强方法,有利于提升语音识别系统的性能。本项目研究丰富了低资源语言语音识别研究,有利于实现藏语的语音识别输入,完善藏语的人机交互,保证网络中传播的藏语信息内容的安全。.除了本项目研究内容,项目组在语音情感识别、印刷藏文吾美文字检测与识别等方面开展了一些研究,取得了一些较好成果。
脱机手写藏文字符识别研究
  • 批准号:
    61462072
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    47.0万元
  • 批准年份:
    2014
  • 负责人:
    黄鹤鸣
  • 依托单位:
藏文字符排序研究
  • 批准号:
    60963016
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    22.0万元
  • 批准年份:
    2009
  • 负责人:
    黄鹤鸣
  • 依托单位:
国内基金
海外基金