课题基金 / 基金详情

Echo - A New Set of High-level Audio Features for Computational Sound Design Systems

Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
Echo - 用于计算声音设计系统的一组新的高级音频功能
批准号:
RGPIN-2021-02893
负责人:
Thorogood, Miles
金额:
$1.75万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
关键词:

项目摘要

项目成果

Thorogood, Miles的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
I propose to advance the knowledge of sound analysis models and algorithms for building new computational tools that can be used to increase the capacity of audio retrieval systems and sound engineering. First, new high-level audio descriptors and an annotated dataset will be investigated. Second, gold-standard predictive and multiple-output neural network models for audio signal recognition will be created. First, from our past work concerned with modeling mood-based audio descriptors, we discovered that many attributes of the human listening experience are missing from the state-of-the-art challenges facing Music Information Retrieval. We design a listening experiment with user and sound design experts to code varying audio stimulus to derive a set of sound describing words. We analyze study results to reveal the correlations between expert and user experience and obtain high-level audio descriptors. From the set of audio descriptors and associated meanings, we will develop a novel lexicon that will address this gap in the literature. Building on our past success of creating large datasets of annotated audio files through crowd-sourcing, we will design an experiment using an online ranking-based questionnaire where annotators make pairwise comparisons between two audio clips based on individual terms from the lexicon. We expect the outcome from this research to result in labelling each audio file in the corpus with a rating value for every concept in the lexicon. This will be the first such dataset annotated with a comprehensive set of sound design descriptors for Music Information Retrieval. Second, To addresses the problem of how a machine can predict the important sonic characteristics for the user experience in video games, VR, and film, we first run a series of machine learning experiments to create gold-standard models for predicting each of the features represented in the lexicon. We will run experiments for training, tuning, and evaluating different supervised machine learning models and assign select audio features and optimized model to each corresponding concept in the lexicon. Next, we investigate a deep neural network model for predicting all (200+) features simultaneously. For this model, we will experiment with different neural network topologies utilizing built-from-scratch techniques and TensorFlow to develop a deep learning model for multi-output regression. The trained model will be evaluated using standard regression metrics and against the set of gold standard models. We plan to build on the Essentia project (an open-source library and tools for audio and music analysis, description and synthesis) to release the gold-standard and multi-output regression DNN models as an open-source library for audio and music analysis. This research directly involves 8 highly qualified personnel working in interdisciplinary teams to become leaders in advanced music information retrieval and creative A.I..
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
  • 批准号:
    DGECR-2021-00050
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2021
  • 负责人:
    Thorogood, Miles
  • 依托单位:
Echo - A New Set of High-level Audio Features for Computational Sound Design Systems
  • 批准号:
    RGPIN-2021-02893
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.75万
  • 财政年份:
    2021
  • 负责人:
    Thorogood, Miles
  • 依托单位:
海外基金