课题基金 / 基金详情

Understanding Speech by Leveraging Both Audio and Lexical Information Channels

Understanding Speech by Leveraging Both Audio and Lexical Information Channels
通过利用音频和词汇信息通道来理解语音
批准号:
2424038
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
已结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
这个项目的首要研究领域是口语处理(SLP),重点是改进从自发的、自然的口语交流中提取信息的方法。该项目的主要目标是利用口语音频通道中编码的信息来改进提取语音语义内容的方法和我们对口语对话的理解。在过去的十年里,NLP研究取得了令人难以置信的进步,然而,有一种媒介尚未如此明显地受益,那就是处理自发的口头对话。这种交流方式可以说是最基本的;人类和他们的语言并肩进化,以优化表达和理解对话的能力。Alexa和Siri等系统的迅速发展,以及对理解人类对话的学术研究,都证明了公众对通过对话语音实现人机交互的巨大兴趣。然而,许多现代NLP已经从几乎完全专注于书面文本发展起来。口语和书面语虽然相关,但本质上是不同的。口语提供了第二种独特的信息传递渠道:音频。语音包含韵律和时间线索,这些线索除了语音的词汇内容外,还传达语义和句法信息。当考察自发的、逐渐产生的、充满不流畅的口头对话时,书面语言和口语之间的差异就会扩大。这些特性为信息传输提供了丰富的带宽,人类可以毫不费力地有效地利用这些带宽,但目前的SLP方法几乎没有利用它们。标准的SLP管道使用级联架构,其中音频仅用于生成文本,下游任务完全基于书面单词。考虑到SLP领域相对年轻的状态,关于有意义的任务和相应的语音表征评估很少达成一致。因此,该项目的第一个组成部分将涉及对目前的方法进行严格审查。第二个部分将侧重于用基于音频的信息增强语音表示。最后,我们将评估哪些信息渠道是重要的,以及何时对机器和人类的理解都是重要的。就目前的研究状况来看,绝大多数的口头互动都无法由机器充分处理。改进对话理解方法可以在人与技术之间提供一个更自然、更容易理解的界面,并解锁以口语传播的大量信息。我对健康应用特别感兴趣。语言是神经健康的最具信息性的可观察探针之一,然而,它在实践中的应用一直受到限制。提高我们对人类语言互动的理解,可以检测出异常的语言,从而更早、更有效地诊断出阿尔茨海默氏症等疾病。
英文摘要
The over-arching research area of this project is spoken language processing (SLP) with a focus on improving methods for extracting information from spontaneous, natural, spoken communication. The key objectives of this project are to make use of information encoded in the audio channel of spoken language to improve both methods to extract the semantic content of speech and our understanding of spoken conversations.The past decade has seen incredible advances in NLP research, however one medium that has yet to benefit so acutely is the processing of spontaneous, spoken dialogue. This modality of communication is arguably the most fundamental; humans and their languages have evolved side-by-side to optimize the ability to express and understand dialogue. There is already massive public interest in enabling human-machine interaction through conversational speech as demonstrated by the burgeoning development of systems like Alexa and Siri, and academic research into understanding human-human conversation. However, much of modern NLP has been developed from an almost-exclusive focus on written text. Though related, spoken and written domains are intrinsically different. Spoken language provides a second, distinct channel of information transfer: audio. The voice contains prosodic and temporal cues which convey semantic and syntactic information alongside the lexical content of speech. The discrepancy between written and spoken language widens when examining spontaneous, spoken dialogue which is generated incrementally and riddled with disfluencies. These features provide a rich bandwidth for information transmission which humans use effectively and with little-to-no effort, however, current SLP methods make little to no use of them. Standard SLP pipelines use a cascading architecture whereby audio is only used to produce a transcript, and downstream tasks are based entirely on the written words. Given that relatively young state of the field of SLP, there is little agreement regarding meaningful tasks and corresponding evaluations of speech representations. Thus, the first component of this project will involve a critical review of current approaches. The second component will focus on augmenting representations of speech with audio-based information. Finally, we will evaluate which information channels are important, and when, for both machine and human understanding. As the state of research currently stands, the vast majority of spoken interaction can't be adequately processed by machines. Improving methods for conversation comprehension could provide a much more natural and accessible interface between people and technology and unlock swathes of information transmitted in spoken language. I am particularly interested in health applications. Speech is one of the most informative observable probes of neurological health, however, its use in practice has been restricted. Improving our understanding of human spoken interaction could enable detection of abnormal speech, allowing diagnosis of diseases like Alzheimer's to be made much earlier and much more efficiently.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金