课题基金 / 基金详情

基于自监督学习的通用语音信号表征学习研究

批准号:
62106140
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
王钰
依托单位:
学科分类:
机器学习
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
王钰

项目摘要

结项摘要

相似基金

相关文献

中文摘要
表征学习对于多媒体信息的处理和理解至关重要, 表征学习算法的性能直接决定了下游的感知和认知任务的性能。近年来,受到语言和视觉领域的启发,无需人工标注信息的自监督学习算法在面向语音信号的表征学习领域开始崭露头角。但是,现有的算法没有充分捕获语音信号的多层次、多尺度、结构化的信息,因此只能应用于有限的下游任务。本项目致力于研究通用的语音信号的自监督表征学习算法,旨在提出能够广泛应用于下游语音信号处理任务的自监督表征学习算法。具体内容包括:研究基于多粒度信息的自监督表征学习和融合算法,捕获语音信号中包含的复杂异构的语义信息;研究基于多任务学习的自监督表征算法,提升语音信号表征对下游任务的泛化能力;研究基于知识蒸馏的模型压缩算法,将模型更高效的应用于实际场景。综上所述,本项目旨在实现高鲁棒性、高泛化能力、低耗的通用语音表征的自监督学习算法,能够为推动智能语音处理技术在智能终端落地提供有力支持。
英文摘要
Learning good representation is a very important task for multi-media signal processing and understanding as it directly impacts the performance of downstream perpetual and cognitive tasks. Recently, inspired its success in text and vision domains, self-supervised learning, which is effective at learning representation without manual labels, is becoming popular in speech signal representation learning. However, existing approaches does not comprehensively take into account the complex hierarchical and multi-scale nature of speech signals. Thus, they can only be applicable on one type of downstream task. This project focuses on self-supervised approaches for universal speech representation learning, the goal of which is to develop models and algorithms for learning general-purpose speech representation that can be useful for a wide range of speech processing tasks. The contribution of this projects includes: In order to capture the complex structured semantic information in speech signals, this project proposes multi-granularity and multi-task self-supervised learning framework; in order to improve the generalization of the representation on the various downstream tasks, this project proposes multi-task self-supervised representation learning algorithms; this project also proposes knowledge distillation-based model compression algorithms, thus making the models suitable for real-world application and products. In summary, this project aims to propose self-supervised learning frameworks that are able to learn universal speech representation with high robustness, good generalization and low operational cost, which can be highly useful in many real-world speech processing-related smart devices.
本项目致力于基于自监督学习的通用语音信号表征学习研究,针对现有算法的挑战与不足,创新性地提出了多层次、多尺度、结构化信息融合的自监督表征学习算法,以及鲁棒异构多自监督学习任务的架构与算法。通过三年的深入研究,项目成功构建了LibriSQA和M3AV两个大规模多模态语音数据集,并开发了轻量化端到端技术框架和基于情感知识的注意力机制,显著提升了多模态情感识别等任务的性能。项目成果在国际顶级会议及期刊上发表论文17篇,申请国家发明专利4项,不仅在学术上产生了广泛影响,还通过技术转化与应用探索,在教育、医疗等领域实现了初步试点应用。同时,项目培养了一批高素质的研究人才,为语音信号处理及相关领域的发展提供了有力支撑。研究成果在智能终端、智能语音交互、智能口语处理、智能医疗等领域展现出广阔的应用前景,为语音及相关序列信号表征的自监督学习理论的完善与算法的实用化做出了重要贡献。
国内基金
海外基金