Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
批准号:
RGPIN-2021-02893
负责人:
Thorogood, Miles
金额:
$1.75万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
I propose to advance the knowledge of sound analysis models and algorithms for building new computational tools that can be used to increase the capacity of audio retrieval systems and sound engineering. First, new high-level audio descriptors and an annotated dataset will be investigated. Second, gold-standard predictive and multiple-output neural network models for audio signal recognition will be created. First, from our past work concerned with modeling mood-based audio descriptors, we discovered that many attributes of the human listening experience are missing from the state-of-the-art challenges facing Music Information Retrieval. We design a listening experiment with user and sound design experts to code varying audio stimulus to derive a set of sound describing words. We analyze study results to reveal the correlations between expert and user experience and obtain high-level audio descriptors. From the set of audio descriptors and associated meanings, we will develop a novel lexicon that will address this gap in the literature. Building on our past success of creating large datasets of annotated audio files through crowd-sourcing, we will design an experiment using an online ranking-based questionnaire where annotators make pairwise comparisons between two audio clips based on individual terms from the lexicon. We expect the outcome from this research to result in labelling each audio file in the corpus with a rating value for every concept in the lexicon. This will be the first such dataset annotated with a comprehensive set of sound design descriptors for Music Information Retrieval. Second, To addresses the problem of how a machine can predict the important sonic characteristics for the user experience in video games, VR, and film, we first run a series of machine learning experiments to create gold-standard models for predicting each of the features represented in the lexicon. We will run experiments for training, tuning, and evaluating different supervised machine learning models and assign select audio features and optimized model to each corresponding concept in the lexicon. Next, we investigate a deep neural network model for predicting all (200+) features simultaneously. For this model, we will experiment with different neural network topologies utilizing built-from-scratch techniques and TensorFlow to develop a deep learning model for multi-output regression. The trained model will be evaluated using standard regression metrics and against the set of gold standard models. We plan to build on the Essentia project (an open-source library and tools for audio and music analysis, description and synthesis) to release the gold-standard and multi-output regression DNN models as an open-source library for audio and music analysis. This research directly involves 8 highly qualified personnel working in interdisciplinary teams to become leaders in advanced music information retrieval and creative A.I..
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
-
批准号:DGECR-2021-00050
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2021
-
负责人:Thorogood, Miles
-
依托单位:
Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
-
批准号:RGPIN-2021-02893
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2021
-
负责人:Thorogood, Miles
-
依托单位:
海外基金