课题基金 / 基金详情

带隐私保护的异步多模态线索语自动识别研究

批准号:
62101351
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
刘李
学科分类:
多媒体信息处理
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
刘李

项目摘要

结项摘要

刘李的其他基金

相似基金

相关文献

中文摘要
线索语(Cued Speech, CS)是一种利用手对音素的提示来辅助唇读的交流方式,主要用于聋哑人之间及其与听力正常人之间的沟通。作为聋哑人的人机互动研究的一部分,CS视频到文本的自动识别可以使得聋哑人的交流更加便利和高效,具有重要的科学意义和应用价值。该领域的研究在近些年已经得到了较好的发展,但仍存在三个难题:CS多模态特征表征能力受限、异步异质多模态特征融合效果不佳,以及缺乏高效且带隐私保护的CS自动识别模型。本项目拟创新性地研究自监督对比学习与无偏学习的协同建模、基于统计分析与生成对抗训练顺序优化的融合机理,以及基于联邦学习的加密共享训练机制,以期实现首个高效且带隐私保护的CS自动识别模型。同时,本项目还将建立首个多语种-多编码者的CS公开数据集,为此研究打好基础。实际上,本项目的研究成果也可服务于听力障碍者早期教育、视听转换以及人机互动等智能应用,促进这些领域的发展。
英文摘要
Cued Speech (CS) is an augmented lip reading system complemented by hand coding in the phonetic level, and it is mainly used for the communication among the deaf people. As one part of research on the human-computer interaction of deaf people, the automatic CS recognition aims to make the communication of deaf people more convenient and efficient, and thus has important social and scientific significance. The automatic CS recognition research based on machine learning has been developed in recent years. However, there are still three challenging problems unsolved that inhibit further development of the CS recognition model: (1) CS multi-modal feature representation ability is still limited; (2) the asynchronous multi-modal feature fusion performance is not satisfactory, and (3) the lack of an efficient and data privacy-protected automatic CS recognition model. In this project, we will try to solve these three problems by studying (1) the collaborative modeling of self-supervised contrastive learning and unbiased learning; (2) statistical analysis and adversarial training-based sequential optimization for multi-modal fusion method, and (3) the encrypted shared training mechanism of the CS recognition federated learning framework. Besides, this project will build the first public “multilingual-multispeaker” CS dataset to lay a solid foundation for this research. In fact, the research outcomes of this project can serve many intelligent applications, such as the early education of the hearing impaired, audio-visual conversion and human-computer interaction, further promoting the development of these fields.
线索语(Cued Speech, CS)是一种利用手对音素的提示来辅助唇读的交流方式,主要用于听障人士之间及其与听力正常人之间的沟通。作为听障人士的人机互动研究的一部分,线索语视频到文本的自动识别可以使得听障人士的交流更加便利和高效,具有重要的科学意义和应用价值。该领域的研究在近些年已经得到了较好的发展,但仍存在多个难题。本项目的主要研究内容就是基于最新的研究工具,以期实现首个高效且带隐私保护的线索语自动识别模型。经过三年的努力,本项目在研究计划的基础上,结合当前最前沿的研究工具和技术,成功实现了预定的各项研究目标。具体成果包括:(a) 提出了联邦线索语识别(FedCSR)框架,在不共享隐私数据的情况下训练线索语自动识别(ACSR)模型。这也是首次在ACSR任务中引入联邦学习。(b) 提出了一种跨模态互学习(Cross-Modal Mutual Learning)框架,用于ACSR任务,解决异步模态(唇形、手势形状、手势位置)融合干扰的问题。此外,我们也构建了首个大规模的普通话多说话人线索语数据集,并在中文、法语、英语上进行实验。(c) 针对ACSR任务中多模态信息融合的挑战,提出了一种高效计算且轻量的多模态融合Transformer模型(EcoCued)。(d) 针对线索语视频生成任务,提出了一种基于扩散模型的Gloss提示细粒度线索语手势生成框架(GlossDiff),以解决现有方法对模板依赖性强、泛化能力弱的问题。项目期间,本项目共发表了25篇高水平学术论文,其中3篇为期刊论文,其他22篇为会议论文,包括IEEE TASLP 1篇、IEEE TPAMI 1篇、NeurIPS 3篇、ACM MM 3篇、ICASSP 6篇、Interspeech 3篇等,完成了项目设定的论文目标。此外,本项目还协助培养了2名博士研究生和2名硕士研究生,达到了人才培养的目标。
个性化与韵律感知的细粒度线索语视频 生成研究
  • 批准号:
    --
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2025
  • 负责人:
    刘李
  • 依托单位:
国内基金
海外基金