Unsupervised Structure Discovery for Semantic Analysis of Audio
Unsupervised Structure Discovery for Semantic Analysis of Audio
复制标题
用于音频语义分析的无监督结构发现
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
B. Raj
中科院分区:
文献类型:
--
作者:
Sourish Chaudhuri;B. Raj
Approaches to audio classification and retrieval tasks largely rely on detection-based discriminative models. We submit that such models make a simplistic assumption in mapping acoustics directly to semantics, whereas the actual process is likely more complex. We present a generative model that maps acoustics in a hierarchical manner to increasingly higher-level semantics. Our model has two layers with the first layer modeling generalized sound units with no clear semantic associations, while the second layer models local patterns over these sound units. We evaluate our model on a large-scale retrieval task from TRECVID 2011, and report significant improvements over standard baselines.