An extensible cluster-graph taxonomy for open set sound scene analysis

An extensible cluster-graph taxonomy for open set sound scene analysis
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Helen L. Bear;Emmanouil Benetos
Helen L. Bear;Emmanouil Benetos
中科院分区:
其他
文献类型:
--
作者:
Helen L. Bear;Emmanouil Benetos

文献摘要

相似文献

我们提出了一种新的可扩展和可划分的分类方法,用于开放场景声音场景分析。这一新模型允许使用有形的描述符和感知标签进行复杂的场景分析。其新颖的结构是簇图,使得每个簇(或子集)可以独立用于诸如办公室声音事件检测之类的目标分析,同时保持整个标签图(超集)上的完整性。设计的主要好处是它的可扩展性,因为在新的数据捕获过程中需要新的标签。此外,使用相同分类的数据集可以很容易地扩展,从而节省了未来的数据收集工作。我们平衡了复杂场景分析所需的细节,并用我们的框架避免了“一切事物的分类”,以确保标签超集中没有重复,并通过DCASE挑战分类进行了演示。
We present a new extensible and divisible taxonomy for open set sound scene analysis. This new model allows complex scene analysis with tangible descriptors and perception labels. Its novel structure is a cluster graph such that each cluster (or subset) can stand alone for targeted analyses such as office sound event detection, whilst maintaining integrity over the whole graph (superset) of labels. The key design benefit is its extensibility as new labels are needed during new data capture. Furthermore, datasets which use the same taxonomy are easily augmented, saving future data collection effort. We balance the details needed for complex scene analysis with avoiding 'the taxonomy of everything' with our framework to ensure no duplicity in the superset of labels and demonstrate this with DCASE challenge classifications.