EAGER: Example-based Audio Editing
EAGER: Example-based Audio Editing
批准号:
1451380
负责人:
Paris Smaragdis
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-05-31
中文摘要
当代技术用户通过编辑照片和视频与之互动,但仍然通过捕获、存储、传输和回放来被动地使用音频。这两种与当代媒体互动的不同方式一直存在,因为目前的软件工具使普通用户很难操纵音频。该项目将开发新的技术,使非专家能够访问音频编辑和操作。这些工具将允许用户通过发声所需编辑、提供所需效果的前后示例、或通过呈现展示所需音频操作的其他录音来指导软件进行音频编辑请求。例如,用户可以向软件发出命令,通过使用低音更响亮的声音或使用中频的鼻音来均衡声音;通过发出“Hello,Hello,Hello……”来模仿所需的效果来增加回声。在较低的音量中使用每个连续的“Hello”;或者通过提供具有所需混响的录音的示例来添加混响。使普通计算机用户更容易操作和编辑音频记录可以影响许多领域,如医学生物声学、地震信号分析、水下监测、音频取证、监视应用、石油勘探探测、对话式数据采集和机械振动测量。这个项目的目标是提供新颖实用的音频工具,让这些领域的非专业从业者能够轻松地完成所需的音频操作。该项目将利用现代信号处理和机器学习技术来产生更直观的界面,帮助人们完成目前困难的音频编辑任务。这将包括开发新的估计器,直接从录音中提取编辑意图参数。该项目将专注于三种不同的编辑操作:均衡、噪音控制和回声/混响。对于每一次手术,都将探索一些不同的方法。例如,对于均衡,一种方法将使用户选择之前和之后的声音以识别他们想要的修改,然后系统将使用频谱去卷积估计来直接计算将之前的声音的频谱映射到之后的声音的频谱的传递函数,并将该函数应用于用户正在编辑的录音。对于噪音控制,一种方法是让用户发出要移除的噪音类型的声音,然后将用户的输入与正在使用低阶谱分解编辑的录音中的相应分量进行匹配。对于混响和回声,一种方法将让用户发出“一,二,三,...”以说明期望的重复次数、时间间隔和回声之间的衰减,然后使用语音检测测量来提取回声参数,同时校正发声错误,例如回声间隔中的随机不一致。该项目将创造关于人工指导和自动音频智能处理如何协同工作的新理论,以解决基本和实际问题。
英文摘要
Contemporary users of technology interact with photos and video by editing them, but still use audio only passively, by capturing, storing, transmitting, and playing it back. These two different ways of interacting with contemporary media persist because current software tools make it very difficult for general users to manipulate audio. This project will develop novel technologies that will make audio editing and manipulation accessible to non-experts. These tools will allow a user to guide the software with audio editing requests by vocalizing the desired edits, providing before/after examples of the desired effects, or by presenting other recordings that exhibit the desired audio manipulations. For example, a user might issue a command to the software to equalize sounds by using a booming voice for more bass, or a nasal tone for middle frequencies; to add echoes by mimicking the desired effect by uttering "hello, hello, hello ..." with each successive "hello" in a lower volume; or to add reverb by providing examples of recordings with the desired reverb. Making it easier for general computer users to manipulate and edit audio recordings can impact many fields, such as medical bioacoustics, seismic signal analysis, underwater monitoring, audio forensics, surveillance applications, oil exploration probing, conversational data gathering, and mechanical vibration measuring. The goals of this project are to provide novel and practical audio tools that will allow non-expert practitioners from these fields to easily achieve required audio manipulations.The project will exploit modern signal processing and machine learning techniques to produce more intuitive interfaces that help people accomplish what are currently difficult audio editing tasks. This will include developing novel estimators to extract editing-intent parameters directly from audio recordings. The project will focus on three different editing operations: equalization, noise control, and echo/reverberation. A number of different approaches will be explored for each operation. For example, for equalization, one approach will have users select before and after sounds to identify their desired modification, and the system will then use spectral deconvolution estimations to directly compute the transfer function that maps the spectrum of the before sound to that of the after sound, and apply that function to the audio recording that the user is editing. For noise control, one approach will have users vocalize what types of noise to remove, and then match the user's input with the corresponding component in the recording that is being edited by using low-rank spectral decomposition. For reverb and echo, one approach will have users utter "one, two, three, ..." to illustrate the desired number of repetitions, temporal spacing, and attenuation between echoes, and then use voice detection measurements to extract the echo parameters, while correcting for vocalization errors such as random inconsistency in the echo spacing. The project will create new theories of how human guidance and automated audio-intelligent processing can work in tandem to solve fundamental and practical problems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Scaling Source Separation to Big Audio Data
-
批准号:1453104
-
项目类别:Continuing Grant
-
资助金额:$54.99万
-
财政年份:2015
-
负责人:Paris Smaragdis
-
依托单位:
III: Small: MicSynth: Enhancing and Reconstructing Sound Scenes from Crowdsourced Recordings
-
批准号:1319708
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2013
-
负责人:Paris Smaragdis
-
依托单位:
海外基金