Iterative Information Fusion in Automatis Speech Recognition According to the Turbo Principle
Iterative Information Fusion in Automatis Speech Recognition According to the Turbo Principle
批准号:
414091002
负责人:
Professor Dr.-Ing. Tim Fingscheidt
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2019
资助国家:
德国
项目状态:
已结题
起止时间:
2018-12-31 至 2022-12-31
中文摘要
信息的智能融合在信息技术的两个对立的大趋势中发挥着主要作用:(1)去中心化(互联网、物联网、去中心化网络控制、传感器网络、工业4.0,.),最近在自动语音识别领域也出现了(2)集中化(Siri,Google Home,Amazon Alexa,YouTube)。这两种趋势的共同点是正在使用多个信息源:它可能是多模式方法(例如,视听语音识别:麦克风、摄像机)或单模态(仅利用麦克风信号的语音识别)。单模式方法可以操作多通道或单通道,在后一种情况下使用例如,在申请人的现有工作中,从数字通信中已知的用于信息的迭代融合的turbo原理已经成功地转移到自动语音识别(ASR)。该项目的一个目标是进一步探索Turbo信息融合在自动语音识别领域中仍然广泛发现的潜力。它不仅能够很好地融合特征表示,而且还可以执行声学模型的融合。由于ASR中的建模同时使用深度神经网络进行,并且各种网络模型拓扑正在研究中,因此融合是一个热门话题,但高性能融合方法很少具有模块化。然而,由于模块化在信息融合的趋势(1)和(2)中几乎是不可或缺的,在本项目中,turbo信息融合将进一步发展成为完全模块化的,从而为广泛的应用提供高度的灵活性和相关性。为什么它的表现这么好?它的性能与信息源的统计依赖性之间的关系如何?利用合成数据进行的受控实验应能提供完美的模型。此外,从数字通信已知的有用的和理论上要求严格的所谓的EXIT图应进一步发展,最终目标是能够预测涡轮信息融合的性能。更重要的是,使用EXIT分析工具,它应该成为可能,融合可以设计的方式,经过几次迭代后,确实获得了高质量的识别结果。最后,我们计划探索ASR与涡轮信息融合与两个以上的信息源或识别器,分别。除了几个互补模型的融合之外,空间分布麦克风和ASR系统的场景也很有趣:turbo信息融合是否能够从空间分布麦克风获得性能增益,例如,一个混响的环境?
英文摘要
The intelligent fusion of information plays a major role in two opposing megatrends of information technology: (1) Decentralization (internet, internet of things, decentralized network control, sensor networks, industry 4.0, ...), and since recently in the field of automatic speech recognition also (2) centralization (Siri, Google Home, Amazon Alexa, YouTube). Both trends have in common that multiple information sources are being used: It may be multimodal approaches (e.g., audiovisual speech recognition: microphone, camera), or uni-modal (speech recognition only with microphone signals). The uni-modal approach may operate multi-channel or single-channel, in the latter case using, e.g., information of two different feature representations.In prior works of the applicant the turbo principle known from Digital Communications for iterative fusion of information has been successfully transferred to automatic speech recognition (ASR). One objective of this project is to further explore the still widely uncovered potential of turbo information fusion in the field of automatic speech recognition. It is not only capable of fusing feature representations very well, but can also perform fusion of acoustic models. Since modelling in ASR is meanwhile performed with deep neural networks, and a variety of network model topologies are subject to research nowadays, fusion is a hot topic, but high-performance fusion approaches rarely come with modularity. However, since modularity in information fusion in both trends (1) and (2) is almost indispensable, in this project turbo information fusion shall be further developed to become completely modular, thereby proving high flexibility and relevance for a wide range of applications.A further objective is to acquire a deeper knowledge of the iteratively operating turbo information fusion. Why is it performing so well? And how about the relation between its performance and statistical dependence of the information sources? Controlled experiments with synthetic data allowing perfect modelling shall provide answers. Also the both useful and theoretically demanding so-called EXIT charts known from Digital Communications shall be developed further with the ultimate goal to be able to predict the performance of turbo information fusion. Even more, using the EXIT analysis tool, it shall become possible that the fusion can be designed in a way such that after a few iterations indeed a high quality recognition result is obtained.Finally, we plan to explore ASR with turbo information fusion with more than two information sources or recognizers, respectively. Besides the fusion of a couple of complementary models the scenario of spatially distributed microphones and ASR systems is of interest: Is turbo information fusion capable of obtaining a performance gain from spatially distributed microphones in, e.g., a reverberant environment?
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bandbreitenerweiterung von Telefonsprachdatenbanken zum Training breitbandiger automatischer Spracherkenner
-
批准号:215637315
-
项目类别:Research Grants (Transfer Project)
-
资助金额:$0.0万
-
财政年份:2012
-
负责人:Professor Dr.-Ing. Tim Fingscheidt
-
依托单位:
Ancient Arabic Document Analysis
-
批准号:142173438
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2009
-
负责人:Professor Dr.-Ing. Tim Fingscheidt
-
依托单位:
Künstliche Erweiterung der Bandbreite von Sprachsignalen mittels phonetischer Transkription
-
批准号:72435333
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2008
-
负责人:Professor Dr.-Ing. Tim Fingscheidt
-
依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
-
批准号:W2433169
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:HAOFEI ZHANG
-
依托单位:
SCIENCE CHINA Information Sciences
-
批准号:61224002
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:宋扉
-
依托单位: