课题基金 / 基金详情

Deep Learning Systems for Musical Audio Generation

Deep Learning Systems for Musical Audio Generation
用于音乐音频生成的深度学习系统
批准号:
RGPIN-2020-05968
负责人:
Oore, Sageev
金额:
$2.11万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Oore, Sageev的其他基金

相似基金

相关文献

中文摘要
翻译
想象一下,您正在为视频创建配乐,并且您有一个机器学习(ML)驱动的音乐创作工具/助手。你的机器学习助手可以生成声音和音乐片段,但有一个问题:它既不能遵循指令,也不能弄清楚需要做什么。对于一个对人类有帮助的助手来说,必须有一种方法来指导它做什么和如何做。 我的研究计划涉及音频和音乐生成领域的ML,而这个建议涉及为音乐和音频注入强大的控制(例如,允许有效的基于ML的工具)。具体来说,我计划在两个主要的音乐环境中探索这一点:(1)生成不同音符的序列(例如钢琴上的按键),以及(2)生成原始音频文件,即声波,一次一个“测量”(其中最低质量的声音每秒至少有16,000个这样的测量)。 由于以下原因,很难精细控制这些领域的发电: (1)我们用来描述声音的词汇是有限的,也是不明确的。我们可以听到这把小提琴听起来比那把“更丰富”,或者这段演讲比那一段“更清晰的节奏”,但我们可能不知道如何量化这些品质。 (2)大多数数据都没有这些标签。 这些困难所隐含的ML挑战是根本性的:在极长序列的大型最小标记数据集中学习底层结构。 那么,我们如何完成控制的任务呢? 我们首先注意到人们是怎么做的:老师告诉学生,“这样演奏”,并提供了一个相关的例子;这个例子立即变得有帮助,因为学生已经有了一个无声的声音心理地图,这个例子成为一个类比,指向地图上的一个点。 我想找到一些技术,以这种方式控制生成模型。这包括:(1)学习良好的地图(~解开潜在的代表性)的声音。例如,沿沿着一个方向移动可能意味着在某种程度上更有节奏。(II)从主要是未标记的数据中学习这些,有效地利用罕见的标记示例(~半监督学习)。 这一点很重要,因为: 1)这个问题所固有的一些ML问题是根本性的,因此它们的解决方案也是根本性的。 2)控制生成模型可以让它们有所帮助。 ..艺术家,因为他们将为创意经济提供有效的创意支持。 ..业余音乐家,因为这些工具可以很容易地有很大的教育价值。 ..对健康有节奏的音乐可以帮助运动康复;想象一下专门为康复设计的不知疲倦和自适应的音乐发生器。精神病诊断有时候是基于非语言的语音质量;想象一下,通过例子来控制语音的生成:“使用像A一样的声音,像B一样的口音,和C一样的韵律。”这样做有助于训练和消除潜在的偏见效应。
英文摘要
Imagine you are creating a soundtrack for a video, and you have a machine learning (ML)-driven music creation tool/assistant. Your ML assistant can generate sounds and musical clips, but there is a problem:it can neither follow instructions, nor figure out what needs to be done. For an assistant to be helpful to humans, there must be a way to direct what it does and how it does it. My research programme is concerned with ML in the sphere of audio and music generation, and this proposal is concerned with imbuing generative models for music and audio with powerful controls (e.g. to allow effective ML-based tools). Specifically, I plan to explore this in two main musical contexts: (1) generating sequences of distinct notes (e.g. keypresses on a piano), and (2) generating raw audio files, i.e. soundwaves, one “measurement” at a time (where the lowest-quality sounds have at least 16,000 such measurements per second). Finely controlling generation in these domains is hard for reasons including: (1) Our vocabulary to describe sounds is limited and ill-defined. We can hear that this violin sounds "richer" than that one, or that this speech has a "more articulated rhythm" than that one, but we may not know how to quantify these qualities. (2) Most data does not come with these labels. The ML challenges implied by these difficulties are fundamental ones: learning underlying structure in large, minimally labelled datasets of extremely long sequences. So how do we approach the task of control? We first notice what people do: a teacher tells a student, "play it this way," providing a related example; that single example becomes immediately helpful, since the student already has a wordless mental map of sounds, and the example becomes an analogy that points to a spot on that map. I want to find techniques to allow control over generative models in such ways. This involves: (I) Learning good maps (~disentangled latent representations) of sound. For example, moving along one direction might mean more rhythmic in some way. (II) Learning these from mainly unlabelled data, making effective use of rare labelled examples (~semi-supervised learning). This is important because: 1) Some ML problems inherent to this problem are fundamental, so their solutions will be fundamental as well. 2) Controlling generative models can allow them to be helpful.. ..to artists, as they will provide effective creativity support to the creative economy. ..to amateur musicians, because such tools can easily have great educational value. ..to health. Rhythmic music can help motor rehabilitation; imagine a tireless and adaptive music generator designed specifically for rehabilitation. Psychiatric diagnoses are sometimes based on non-verbal speech qualities; imagine controlling speech generation by examples: “use a voice like Person A, with an accent like Person B, and with the prosody of Person C.” and in doing so, helping training and removing potential bias effects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Controlling Generative Musical Systems: Getting the Right Data & Using the Right Instrument
  • 批准号:
    RTI-2023-00594
  • 项目类别:
    Research Tools and Instruments
  • 资助金额:
    $4.94万
  • 财政年份:
    2022
  • 负责人:
    Oore, Sageev
  • 依托单位:
Deep Learning Systems for Musical Audio Generation
  • 批准号:
    RGPIN-2020-05968
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.11万
  • 财政年份:
    2022
  • 负责人:
    Oore, Sageev
  • 依托单位:
Deep Learning Systems for Musical Audio Generation
  • 批准号:
    RGPIN-2020-05968
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.11万
  • 财政年份:
    2021
  • 负责人:
    Oore, Sageev
  • 依托单位:
Adaptive high degree-of-freedom interaction techniques
  • 批准号:
    298224-2007
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2013
  • 负责人:
    Oore, Sageev
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: