Unsupervised Structure Discovery for Semantic Analysis of Audio

Unsupervised Structure Discovery for Semantic Analysis of Audio
复制标题

用于音频语义分析的无监督结构发现

DOI:
--
复制
发表时间:
2012
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
B. Raj
B. Raj
中科院分区:
--
文献类型:
--
作者:
Sourish Chaudhuri;B. Raj

文献摘要

被引文献

相似文献

音频分类和检索任务的方法在很大程度上依赖于基于检测的判别模型。我们认为,这种模型在将声学直接映射到语义方面做了一个简单的假设,而实际的过程可能更复杂。我们提出了一个生成模型,映射声学在一个层次的方式越来越高层次的语义。我们的模型有两层,第一层建模广义的声音单元没有明确的语义关联,而第二层模型的本地模式在这些声音单元。我们评估了我们的模型在TRECVID 2011的大规模检索任务,并报告了标准基线的显着改进。
Approaches to audio classification and retrieval tasks largely rely on detection-based discriminative models. We submit that such models make a simplistic assumption in mapping acoustics directly to semantics, whereas the actual process is likely more complex. We present a generative model that maps acoustics in a hierarchical manner to increasingly higher-level semantics. Our model has two layers with the first layer modeling generalized sound units with no clear semantic associations, while the second layer models local patterns over these sound units. We evaluate our model on a large-scale retrieval task from TRECVID 2011, and report significant improvements over standard baselines.