Deep Learning Systems for Musical Audio Generation
Deep Learning Systems for Musical Audio Generation
批准号:
RGPIN-2020-05968
负责人:
Oore, Sageev
金额:
$2.11万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31
中文摘要
假设您正在为视频创建配乐,并且您有一个机器学习(ML)驱动的音乐创作工具/助手。您的ML助手可以生成声音和音乐片段,但有一个问题:它既不能遵循指令,也不能计算出需要做什么。对于一个对人类有帮助的助手来说,必须有一种方法来指导它做什么以及如何做。
我的研究计划是关于音频和音乐生成领域中的ML,这个建议是关于用强大的控制来灌输音乐和音频的生成模型(例如,允许有效的基于ML的工具)。具体地说,我计划在两个主要音乐环境中探索这一点:(1)生成不同音符的序列(例如钢琴上的按键),以及(2)生成原始音频文件,即声波,一次一个“测量”(其中最低质量的声音每秒至少有16,000个这样的测量)。
在这些领域很难精确控制发电,原因包括:
(1)我们描述声音的词汇量有限,定义不清。我们可以听到这把小提琴听起来比那把“丰富”,或者这一次的演讲比那一次“节奏更清晰”,但我们可能不知道如何量化这些品质。
(2)大多数数据没有这些标签。
这些困难所隐含的ML挑战是根本性的:在极长序列的大型、最小标签数据集中学习底层结构。
那么,我们如何处理控制的任务呢?我们首先注意到人们是怎么做的:一位老师告诉学生,“这样玩”,提供了一个相关的例子;这个例子立即变得有帮助,因为学生已经有了一个无词的声音心理地图,这个例子变成了一个指向地图上一个点的类比。
我想找到一些技术,允许以这种方式控制生成性模型。这包括:(I)学习好的声音映射(~去纠缠的潜在表示)。例如,沿着一个方向移动可能意味着在某种程度上更有节奏。(2)主要从未标记的数据中学习这些数据,有效利用罕见的标记示例(~半监督学习)。
这一点很重要,因为:
1)这个问题所固有的一些ML问题是根本性的,因此其解决方案也将是根本性的。
2)控制生成性模型可以让它们发挥作用。
..对艺术家来说,因为他们将为创意经济提供有效的创意支持。
..对于业余音乐家来说,因为这样的工具很容易就有很大的教育价值。
..为了健康。有节奏的音乐可以帮助运动康复;想象一下专门为康复设计的不知疲倦和自适应的音乐发生器。精神病学的诊断有时是基于非语言的语音质量;想象一下通过例子来控制语音的产生:“使用像A人一样的声音,像B人一样的口音,用C人的韵律。”在这样做的过程中,帮助培训和消除潜在的偏见影响。
英文摘要
Imagine you are creating a soundtrack for a video, and you have a machine learning (ML)-driven music creation tool/assistant. Your ML assistant can generate sounds and musical clips, but there is a problem:it can neither follow instructions, nor figure out what needs to be done. For an assistant to be helpful to humans, there must be a way to direct what it does and how it does it.
My research programme is concerned with ML in the sphere of audio and music generation, and this proposal is concerned with imbuing generative models for music and audio with powerful controls (e.g. to allow effective ML-based tools). Specifically, I plan to explore this in two main musical contexts: (1) generating sequences of distinct notes (e.g. keypresses on a piano), and (2) generating raw audio files, i.e. soundwaves, one “measurement” at a time (where the lowest-quality sounds have at least 16,000 such measurements per second).
Finely controlling generation in these domains is hard for reasons including:
(1) Our vocabulary to describe sounds is limited and ill-defined. We can hear that this violin sounds "richer" than that one, or that this speech has a "more articulated rhythm" than that one, but we may not know how to quantify these qualities.
(2) Most data does not come with these labels.
The ML challenges implied by these difficulties are fundamental ones: learning underlying structure in large, minimally labelled datasets of extremely long sequences.
So how do we approach the task of control? We first notice what people do: a teacher tells a student, "play it this way," providing a related example; that single example becomes immediately helpful, since the student already has a wordless mental map of sounds, and the example becomes an analogy that points to a spot on that map.
I want to find techniques to allow control over generative models in such ways. This involves: (I) Learning good maps (~disentangled latent representations) of sound. For example, moving along one direction might mean more rhythmic in some way. (II) Learning these from mainly unlabelled data, making effective use of rare labelled examples (~semi-supervised learning).
This is important because:
1) Some ML problems inherent to this problem are fundamental, so their solutions will be fundamental as well.
2) Controlling generative models can allow them to be helpful..
..to artists, as they will provide effective creativity support to the creative economy.
..to amateur musicians, because such tools can easily have great educational value.
..to health. Rhythmic music can help motor rehabilitation; imagine a tireless and adaptive music generator designed specifically for rehabilitation. Psychiatric diagnoses are sometimes based on non-verbal speech qualities; imagine controlling speech generation by examples: “use a voice like Person A, with an accent like Person B, and with the prosody of Person C.” and in doing so, helping training and removing potential bias effects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Controlling Generative Musical Systems: Getting the Right Data & Using the Right Instrument
-
批准号:RTI-2023-00594
-
项目类别:Research Tools and Instruments
-
资助金额:$4.94万
-
财政年份:2022
-
负责人:Oore, Sageev
-
依托单位:
Deep Learning Systems for Musical Audio Generation
-
批准号:RGPIN-2020-05968
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2022
-
负责人:Oore, Sageev
-
依托单位:
Deep Learning Systems for Musical Audio Generation
-
批准号:RGPIN-2020-05968
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2021
-
负责人:Oore, Sageev
-
依托单位:
Adaptive high degree-of-freedom interaction techniques
-
批准号:298224-2007
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2013
-
负责人:Oore, Sageev
-
依托单位:
Adaptive high degree-of-freedom interaction techniques
-
批准号:298224-2007
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2010
-
负责人:Oore, Sageev
-
依托单位:
Adaptive high degree-of-freedom interaction techniques
-
批准号:298224-2007
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2009
-
负责人:Oore, Sageev
-
依托单位:
Adaptive high degree-of-freedom interaction techniques
-
批准号:298224-2007
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2008
-
负责人:Oore, Sageev
-
依托单位:
Adaptive high degree-of-freedom interaction techniques
-
批准号:298224-2007
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2007
-
负责人:Oore, Sageev
-
依托单位:
Interactive tools for computer animation
-
批准号:298224-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2006
-
负责人:Oore, Sageev
-
依托单位:
Interactive tools for computer animation
-
批准号:298224-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2005
-
负责人:Oore, Sageev
-
依托单位:
Interactive tools for computer animation
-
批准号:298224-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2004
-
负责人:Oore, Sageev
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于有向超图的大型个性化e-learning学习过程模型的自动生成与优化
-
批准号:61572533
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:孙雪冬
-
依托单位:
E-Learning中学习者情感补偿方法的研究
-
批准号:61402392
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2014
-
负责人:秦继伟
-
依托单位: