Self-supervised graph-based representation for language and speaker detection
Self-supervised graph-based representation for language and speaker detection
批准号:
21K17776
负责人:
沈 鵬
金额:
$2.91万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Early-Career Scientists
财政年份:
2021
资助国家:
日本
项目状态:
已结题
起止时间:
2021-04-01 至 2024-03-31
中文摘要
我专注于研究如何更好地表示语言识别和语音识别任务的语音信号。具体来说,为推进本项目,主要做了以下工作:1.项目进度;改进语言识别(LID)语音信号的表示:我们提出了一种新的基于换能器的语言嵌入方法,通过将RNN换能器模型集成到语言嵌入框架中。利用RNN换能器语言表征能力的优势,该方法可以同时利用语音感知声学特征和显式语言特征来完成LID任务。该研究论文被Interspeech 2022收录。此外,我们在NICT LID系统上进一步研究了这些技术,也证明了跨通道数据的鲁棒性。另一项工作着重于改进汉语ASR的RNN-T。我建议使用一种新颖的发音感知独特的字符编码来构建端到端的基于rnn的普通话ASR系统。提出的编码是基于发音的音节和字符索引(CI)的组合。通过引入CI, RNN-T模型在利用语音信息提取建模单元的同时克服了同音问题。利用所提出的编码方法,可以通过一对一映射将模型输出转换为最终识别结果。论文已被IEEE SLT 2022接受。
英文摘要
I focused on investigating how to better represent speech signals for both language recognition and speech recognition tasks. In detail, the following work was done to progress this project:1. Improving the representation of speech signal for language identification (LID): We propose a novel transducer-based language embedding approach for LID tasks by integrating an RNN transducer model into a language embedding framework. Benefiting from the advantages of the RNN transducer's linguistic representation capability, the proposed method can exploit both phonetically-aware acoustic features and explicit linguistic features for LID tasks. The research paper was accepted by Interspeech 2022. Additionally, we further investigated these techniques on the NICT LID system, which also demonstrated robustness on cross-channel data.2. Another work focuses on improving RNN-T for Mandarin ASR. I propose to use a novel pronunciation-aware unique character encoding for building end-to-end RNN-T-based Mandarin ASR systems. The proposed encoding is a combination of pronunciation-based syllable and character index (CI). By introducing the CI, the RNN-T model can overcome the homophone problem while utilizing the pronunciation information for extracting modeling units. With the proposed encoding, the model outputs can be converted into the final recognition result through a one-to-one mapping. This paper was accepted by IEEE SLT 2022.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2204.03888
发表时间:
2022-04
期刊:
影响因子:
--
作者:
[Peng Shen;Xugang Lu;H. Kawai]
通讯作者:
Peng Shen;Xugang Lu;H. Kawai
DOI:
10.48550/arxiv.2203.17036
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
作者:
[Xugang Lu;Peng Shen;Yu Tsao;H. Kawai]
通讯作者:
Xugang Lu;Peng Shen;Yu Tsao;H. Kawai
Siamese Neural Network with Joint Bayesian Model Structure for Speaker Verification
用于说话人验证的联合贝叶斯模型结构的连体神经网络
DOI:
--
发表时间:
2021
期刊:
影响因子:
--
作者:
[Lu Xugang, Shen Peng, Tsao Yu, Kawai Hisashi]
通讯作者:
Kawai Hisashi
DOI:
10.1109/taslp.2021.3129360
发表时间:
2021-01
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[Xugang Lu;Peng Shen;Yu-Yu Tsao-Yu;H. Kawai]
通讯作者:
Xugang Lu;Peng Shen;Yu-Yu Tsao-Yu;H. Kawai
An investigation of generative acoustic latent representations for meeting speech recognition and summarization
-
批准号:24K15004
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$3.0万
-
财政年份:2024
-
负责人:沈 鵬
-
依托单位:
海外基金