Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
批准号:
RGPIN-2021-02893
负责人:
Thorogood, Miles
金额:
$1.75万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
我建议提高声音分析模型和算法的知识,以建立新的计算工具,这些工具可以用于增加音频检索系统和声音工程的能力。首先,将研究新的高级音频描述符和带注释的数据集。其次,将建立用于音频信号识别的黄金标准预测和多输出神经网络模型。首先,从我们过去对基于情绪的音频描述符建模的工作中,我们发现音乐信息检索面临的最新挑战中缺失了人类听力体验的许多属性。我们与用户和声音设计专家一起设计了一个听力实验,对不同的音频刺激进行编码,以得出一组声音描述词。我们对研究结果进行分析,以揭示专家体验和用户体验之间的相关性,并获得高级别的音频描述符。从音频描述符和相关含义的集合中,我们将开发一个新的词典,以解决文学中的这一差距。在我们过去通过众包创建带注释的音频文件的大型数据集的成功基础上,我们将使用基于在线排名的调查问卷设计一个实验,注释员根据词典中的单个术语对两个音频片段进行配对比较。我们预计这项研究的结果将导致为语料库中的每个音频文件标记词典中每个概念的评级值。这将是第一个这样的数据集,用一套全面的音乐信息检索声音设计描述符进行注释。其次,为了解决机器如何预测视频游戏、VR和电影中用户体验的重要声音特征的问题,我们首先运行了一系列机器学习实验,以创建黄金标准模型来预测词典中表示的每个特征。我们将进行实验,对不同的有监督机器学习模型进行训练、调整和评估,并将选定的音频特征和优化模型分配给词典中每个相应的概念。接下来,我们研究了一种同时预测所有(200+)个特征的深度神经网络模型。对于这个模型,我们将使用从头开始的技术和TensorFlow来试验不同的神经网络拓扑,以开发用于多输出回归的深度学习模型。训练后的模型将使用标准回归度量并对照黄金标准模型集进行评估。我们计划在Essentia项目(音频和音乐分析、描述和合成的开源库和工具)的基础上发布黄金标准和多输出回归DNN模型,作为音频和音乐分析的开源库。这项研究直接涉及8名在跨学科团队中工作的高素质人员,他们将成为高级音乐信息检索和创造性人工智能领域的领导者。
英文摘要
I propose to advance the knowledge of sound analysis models and algorithms for building new computational tools that can be used to increase the capacity of audio retrieval systems and sound engineering. First, new high-level audio descriptors and an annotated dataset will be investigated. Second, gold-standard predictive and multiple-output neural network models for audio signal recognition will be created. First, from our past work concerned with modeling mood-based audio descriptors, we discovered that many attributes of the human listening experience are missing from the state-of-the-art challenges facing Music Information Retrieval. We design a listening experiment with user and sound design experts to code varying audio stimulus to derive a set of sound describing words. We analyze study results to reveal the correlations between expert and user experience and obtain high-level audio descriptors. From the set of audio descriptors and associated meanings, we will develop a novel lexicon that will address this gap in the literature. Building on our past success of creating large datasets of annotated audio files through crowd-sourcing, we will design an experiment using an online ranking-based questionnaire where annotators make pairwise comparisons between two audio clips based on individual terms from the lexicon. We expect the outcome from this research to result in labelling each audio file in the corpus with a rating value for every concept in the lexicon. This will be the first such dataset annotated with a comprehensive set of sound design descriptors for Music Information Retrieval. Second, To addresses the problem of how a machine can predict the important sonic characteristics for the user experience in video games, VR, and film, we first run a series of machine learning experiments to create gold-standard models for predicting each of the features represented in the lexicon. We will run experiments for training, tuning, and evaluating different supervised machine learning models and assign select audio features and optimized model to each corresponding concept in the lexicon. Next, we investigate a deep neural network model for predicting all (200+) features simultaneously. For this model, we will experiment with different neural network topologies utilizing built-from-scratch techniques and TensorFlow to develop a deep learning model for multi-output regression. The trained model will be evaluated using standard regression metrics and against the set of gold standard models. We plan to build on the Essentia project (an open-source library and tools for audio and music analysis, description and synthesis) to release the gold-standard and multi-output regression DNN models as an open-source library for audio and music analysis. This research directly involves 8 highly qualified personnel working in interdisciplinary teams to become leaders in advanced music information retrieval and creative A.I..
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
-
批准号:DGECR-2021-00050
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2021
-
负责人:Thorogood, Miles
-
依托单位:
Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
-
批准号:RGPIN-2021-02893
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2021
-
负责人:Thorogood, Miles
-
依托单位:
海外基金