CAREER: Scaling Source Separation to Big Audio Data
CAREER: Scaling Source Separation to Big Audio Data
批准号:
1453104
负责人:
Paris Smaragdis
金额:
$54.99万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-08-01 至 2020-07-31
中文摘要
我们生活的世界是由混合信号组成的。几乎不可能轻易地获得语音、音乐、环境、机械、水下或生物医学声音的干净录音。这进而由于不想要的元素(例如,语音识别中的背景噪声)的存在而使任何进一步的处理复杂化。传统上,我们通过使用源分离和去噪方法来解决这个问题,这些方法允许我们只从混合物中提取所需的信号。不幸的是,这些方法不能扩展到大数据?这意味着我们今天收集的大部分音频数据仍然处于不适合自动内容分析或进一步处理的状态。这个项目解决了在面对非常大的数据集时使用源分离方法的问题。它考虑了使用现代数据分析方法来有效地处理大量数据,并考虑了训练对大信号语料库的影响,以提高源分离性能。最终目标是通过利用大量信号集合来提高源分离性能,并通过提高现代源分离算法的效率来实现大规模信号分析。为了能够在如此大的规模上进行处理,该项目利用了流形结构分析、频谱分解的压缩方法、散列策略和量化的最新发展。通过使用量化流形表示,可以使用紧凑且高效可访问的结构来近似大信号数据集。这样的表示然后可以用作源分离算法的先验,并帮助引导它提取与它们匹配的信号。通过使用通货紧缩方法,在大数据上应用这种模型来执行源分离的速度进一步加快。这个项目不是执行教科书(和计算密集型)的模型匹配过程,而是使用贪婪的方法来快速提取目标组件,同时绕过许多不必要的计算。最后,鉴于关于统一多个源分离模型的最新研究,该项目考虑将这些概念同时应用于多个数据分析模型(例如HMM模型、连续动态系统等),从而变得与广泛的信号和混合情况相关。
英文摘要
The world we live in is composed out of mixed signals. It is practically impossible to easily obtain a clean recording of speech, music, environmental, mechanical, underwater, or biomedical sounds. This, in turn, complicates any further processing due to the presence of unwanted elements (e.g. background noise in speech recognition). We traditionally address this issue by using source separation and denoising methods that allow us to extract only a desired signal from a mixture. Unfortunately such methods do not scale at ?big data? levels, which means that most of the audio data we gather today remains at a state unsuitable for automatic content analysis or further processing. This project addresses the use of source separation methods when confronted with very large data sets. It considers use of modern data analysis methods to efficiently process large amounts of data, and also the effects of training on large signal corpora in order to improve source separation performance. The ultimate goals are to improve source separation performance by leveraging large signal collections, and to enable large-scale signal analysis by making modern source separation algorithms more efficient.In order to enable processing at such large scales, this project takes advantage of recent developments in manifold structure analysis, deflation methods for spectral decompositions, hashing strategies, and quantization. Using a quantized manifold representation, large signal data sets can be approximated using a compact and efficiently accessible structure. Such representations can then be used as priors to a source separation algorithm and help guide it to extract signals that match them. Applying such models on large data to perform source separation is further accelerated by making use of a deflation method. Instead of performing the textbook (and computationally intensive) model matching process, this project uses a greedy approach to quickly extract target components while bypassing many unnecessary calculations. Finally, given the latest research on unifying multiple models of source separation, this project considers the application of such concepts to multiple data analysis models at once (such as HMM models, continuous dynamical systems, etc.), thereby becoming relevant to a wide range of signals and mixing situations.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Example-based Audio Editing
-
批准号:1451380
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2014
-
负责人:Paris Smaragdis
-
依托单位:
III: Small: MicSynth: Enhancing and Reconstructing Sound Scenes from Crowdsourced Recordings
-
批准号:1319708
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2013
-
负责人:Paris Smaragdis
-
依托单位:
海外基金