课题基金 / 基金详情

SHARES - System-on-chip Heterogeneous Architecture Recognition Engine for Speech

SHARES - System-on-chip Heterogeneous Architecture Recognition Engine for Speech
SHARES - 用于语音的片上系统异构架构识别引擎
批准号:
EP/D048605/1
负责人:
Roger Woods
金额:
$64.15万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2006
资助国家:
英国
项目状态:
已结题
起止时间:
2006 至 --

项目摘要

项目成果

Roger Woods的其他基金

相似基金

相关文献

中文摘要
翻译
可行的、健壮的语音识别系统的可用性有可能彻底改变人们与移动的技术交互的方式。这意味着超越简单的呼叫总部类型命令,能够向您的移动终端口述任意、大量的电子邮件,并使用自然语音可靠、高效地访问其日益复杂的功能。这将在许多重要的应用场景中为最广泛的潜在用户释放下一代便携式技术的潜力,例如紧急服务和军事环境以及时间效率高的商业和消费者使用。然而,当前的问题是,满足用户对自然性和鲁棒性的期望所需的算法复杂性的增加远远超过了当前嵌入式处理器技术的处理和功率能力预测。因此,需要新的体系结构来从根本上推进移动的和嵌入式设备的最先进的识别技术的步伐。用于移动的应用的商业语音识别引擎通常是桌面解决方案的小尺寸版本,具有可接受质量的识别功能高度受限于任何给定嵌入式平台上可用的处理和功率预算。应用程序通常受限于一些命令和名称或歌曲列表。相比之下,最先进的自然无约束语音研究系统在2.8 GHz Xeon处理器上的实时运行速度要慢200倍。此外,算法研究,以保持识别精度在嘈杂的操作环境中,被认为是必要的广泛采用的识别技术,指向更大的复杂性。因此,算法要求与传统处理器平台的处理和功率能力之间的差距甚至进一步增长。对于大词汇量连续语音识别(LVCSR)引擎,解码最可能的单词序列本质上是在所有可能的单词组合上的极大规模搜索问题。为了科普巨大的潜在搜索空间,在解码过程中动态创建的搜索网络,直到最近,被认为是唯一可行的方法来实现大词汇识别。静态网络太大了,除了更受限制的词汇任务。然而,在一个显着偏离公认的智慧,充分扩展的大词汇量的静态搜索网络之前,解码已被重要地证明使用加权有限状态转换器(WFST)。WFST结构创造了相当大的潜力,实现高效的正则化解码架构,我们打算利用。据我们所知,我们将是第一个专门利用加权有限状态传感器网络解码框架在新的硬件架构低功耗大复杂度语音识别。
英文摘要
The availability of viable, robust speech recognition systems has the potential to revolutionalise the way that people interact with mobile technology. This implies moving beyond simple call home type commands, to being able to dictate arbitrary, extensive e-mails to your mobile device and to reliably and efficiently access its increasingly complex features using natural speech. This will unlock the potential of next generation portable technology to the widest range of potential users in many important application scenarios e.g. for emergency services and military environments as well as time-efficient business and consumer usage. The current issue is, however, that the increasing algorithmic complexity needed to meet user expectations for naturalness and robustness far exceeds the processing and power capabilities forecast for current embedded processor technology. New architectures are therefore needed to radically advance the pace of state-of-the-art recognition technology for mobile and embedded devices.Commercial speech recognition engines for mobile applications are typically small footprint versions of desktop solutions, with the recognition functionality for acceptable quality highly constrained to the processing and power budget available on any given embedded platform. Applications are typically constrained to a few commands and name or song lists. In comparison, state-of-the art research systems on natural unconstrained speech run up to 200-times slower than real-time on 2.8 GHz Xeon processors. In addition, algorithmic research to maintain recognition accuracy in acoustically noisy operating environments, considered essential to widespread adoption of recognition technology, points towards even greater complexity. The gap between algorithmic requirements and the processing and power capability of conventional processor platforms is thus growing even further.For large vocabulary continuous speech recognition (LVCSR) engines, decoding the most likely sequence of words is essentially an extremely large scale search problem over all possible word combinations. To cope with the huge size of the potential search space, search networks created dynamically during decoding were, until recently, considered the only viable approach to realise large vocabulary recognition. Static networks were too big for all but more constrained vocabulary tasks. However, in a significant departure from accepted wisdom, full expansion of large vocabulary static search networks prior to decoding has been importantly demonstrated using the Weighted Finite State Transducer (WFST). The WFST structure creates considerable potential for achieving efficient regularised decoding architectures, which we intend to exploit. To our knowledge, we would be the first to specifically exploit the Weighted Finite State Transducer network decoding framework in novel hardware architectures for low power large complexity speech recognition.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
FPGA Implementation of a Pipelined Gaussian Calculation for HMM-Based Large Vocabulary Speech Recognition
基于 HMM 的大词汇量语音识别的流水线高斯计算的 FPGA 实现
DOI: 10.1155/2011/697080
发表时间: 2011
期刊: International Journal of Reconfigurable Computing
影响因子: 4.3
作者: [Veitch R]
通讯作者: Veitch R
Noise Compensation and Missing-Feature Decoding for Large Vocabulary Speech Recognition in Noise
噪声中大词汇量语音识别的噪声补偿和缺失特征解码
DOI: --
发表时间: 2008
期刊: INTERSPEECH
影响因子: --
作者: [Lv, J]
通讯作者: Lv, J
Replacing Uncertainty Decoding with Subband Re-estimation for Large Vocabulary Speech Recognition in Noise
用子带重估计代替不确定性解码,实现噪声中的大词汇量语音识别
DOI: --
发表时间: 2009
期刊: INTERSPEECH
影响因子: --
作者: [Lv, J]
通讯作者: Lv, J
eFutures: Electronic systems technology for emerging challenges
  • 批准号:
    EP/X039218/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $96.9万
  • 财政年份:
    2023
  • 负责人:
    Roger Woods
  • 依托单位:
RAPID: ReAl-time Process ModellIng and Diagnostics: Powering Digital Factories
  • 批准号:
    EP/V02860X/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $51.43万
  • 财政年份:
    2022
  • 负责人:
    Roger Woods
  • 依托单位:
eFutures 2.0: Addressing Future Challenges
  • 批准号:
    EP/S032045/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $61.68万
  • 财政年份:
    2019
  • 负责人:
    Roger Woods
  • 依托单位:
Kelvin-2
  • 批准号:
    EP/T022175/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $420.99万
  • 财政年份:
    2019
  • 负责人:
    Roger Woods
  • 依托单位:
国内基金
海外基金
基于铁死亡探讨黄芪甲苷调控System/Xc-/GSH/GPX4信号通路在神经损伤性勃起功能障碍治疗中的作用及机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    马轲
  • 依托单位:
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
TBX1/LKB1轴阻断system Xc活性调控AML细胞铁死亡的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    15.0万元
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
TET2通过调控BAP1-System Xc-轴促进紫拉非尼诱导的肝细胞癌铁死亡的机制研究
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    --
  • 依托单位: