Expressive data augmentation in deep learning
Expressive data augmentation in deep learning
批准号:
RGPIN-2022-04651
负责人:
Summers, Cecilia
金额:
$1.82万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
深度学习是机器学习的一个子领域,机器学习是人工智能的一种,其目标是自动学习如何使用数据解决问题。例如,深度学习中的一项典型任务称为“图像分类”,它包括学习如何在给定带有相应类别的图像数据集时将图像分类为不同的类别(例如,“猫”、“狗”、“人”)。对于人类来说,这项任务很容易,但由于计算机只能将一幅图像“看到”为一串1和0,因此很难将如何解决问题编码到算法中,这是一套计算机可以遵循的机械指令。近年来,深度学习在设计这样的算法很困难的地方实现了新的应用,在涉及图像、音频和语言等各种任务方面取得了进展。深度学习的一个关键限制是,它通常需要大量数据才能很好地工作。例如,在图像分类中,需要几千到几百万个标记图像才能获得合理的性能,这对于大多数应用来说是一种令人望而却步的成本。为了帮助补偿,从现有数据人工生成新数据是很常见的,这一过程被称为“数据增强”。一个基本的例子是随机地对图像的亮度进行轻微的更改,同时保持其标签-猫的图像仍然是猫的图像,即使亮度发生了很小的变化。数据扩充具有扩大用于学习算法的数据集的有效大小的效果,而不需要昂贵的新数据收集。尽管它有很大的实用性,但在使用数据增强时存在一些挑战,这是我的研究想要解决的问题。例如,当将其应用于新问题时,需要定义其基本运算(例如随机亮度变化本身)并确定其精确强度,这可能是代价高昂的。我的研究将通过学习在保留所需标签的同时改变图像精确外观的操作来自动定义增强操作。然后,为了调整每个操作的强度,我的研究将调查在没有数据扩充的情况下训练的算法的学习行为;如果算法输出相对于特定操作有很大差异,那么使用大量的算法作为数据扩充很可能会提高算法对它的稳健性。如果成功,我的研究将允许自动创建表现力数据增强策略,大大减少在加拿大整个研究和行业中解锁深度学习新应用所需的数据量。这一技术的一个特别令人兴奋的应用是在医学上,因为大多数医疗任务可用的数据是有限的。理想情况下,改进的和新的诊断方法的发展是可能的,从而促进加拿大的医学研究,并最终促进加拿大公众的整体健康。
英文摘要
Deep learning is a subfield of machine learning, a type of artificial intelligence, whose goal is to automatically learn how to solve problems using data. For example, a typical task in deep learning is called "image classification", and consists of learning how to categorize images into different categories (e.g. "cat", "dog", "human") when given a dataset of images labeled with their corresponding category. For a human, this task is easy, but since computers can only "see" an image as a bunch of ones and zeros, it is hard to encode how to solve the problem into an algorithm, a set of mechanical instructions that a computer can follow. In recent years, deep learning has enabled new applications where designing such algorithms is difficult, making advances on tasks involving images, audio, and language, among a variety of others. One key limitation of deep learning is that it typically requires a large amount of data in order to work well. In image classification, for example, several thousand to several million labeled images are required for reasonable performance, a prohibitive cost for most applications. To help compensate, it is common to artificially generate new data from existing data, a process known as "data augmentation". A basic example of this is to randomly make slight alterations to an image's brightness while maintaining its label - an image of a cat is still an image of a cat, even if the brightness is changed by a small amount. Data augmentation has the effect of expanding the effective size of the dataset used to learn algorithms without requiring the costly collection of new data. Despite its large utility, a number of challenges exist when using data augmentation, which my research intends to solve. When applying it to a new problem, for example, one needs to define its basic operations (e.g. the random brightness change itself) and decide on their precise strengths, which may be costly. My research will define augmentation operations automatically by learning operations that vary the precise appearance of images while preserving their desired labels. Then, to tune the strength of each operation, my research will investigate the learned behavior of algorithms trained without data augmentation; if algorithm output varies greatly with respect to a particular operation, then it is likely that using strong amounts of it as data augmentation will improve an algorithm's robustness to it. If successful, my research will allow for the automatic creation of expressive data augmentation policies, substantially reducing the amount of data required to unlock new applications of deep learning throughout both research and industry in Canada. One particularly exciting application of this is in medicine, since the data available for most medical tasks is limited. Ideally, the development of both improved and novel diagnostics may be possible, advancing Canadian medical research and eventually the health of the Canadian public as a whole.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Expressive data augmentation in deep learning
-
批准号:DGECR-2022-00408
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2022
-
负责人:Summers, Cecilia
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
复杂数据下半参数转换模型及其在老年慢性病发展中的应用研究
-
批准号:72101261
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:孙韬
-
依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
-
批准号:--
-
项目类别:--
-
资助金额:40万元
-
批准年份:2020
-
负责人:Vikrant Gupta
-
依托单位:
基于高频信息下高维波动率矩阵估计及应用
-
批准号:71901118
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2019
-
负责人:穆燕
-
依托单位:
半参数空间自回归面板模型的有效估计与应用研究
-
批准号:71961011
-
项目类别:地区科学基金项目
-
资助金额:16.0万元
-
批准年份:2019
-
负责人:丁飞鹏
-
依托单位:
高频数据波动率统计推断、预测与应用
-
批准号:71971118
-
项目类别:面上项目
-
资助金额:50.0万元
-
批准年份:2019
-
负责人:孔新兵
-
依托单位:
经济管理中复杂数据和复杂行为的分析方法及其应用
-
批准号:71931004
-
项目类别:重点项目
-
资助金额:230.0万元
-
批准年份:2019
-
负责人:周勇
-
依托单位:
基于个体分析的投影式非线性非负张量分解在高维非结构化数据模式分析中的研究
-
批准号:61502059
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2015
-
负责人:刘昶
-
依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
-
批准号:61373035
-
项目类别:面上项目
-
资助金额:77.0万元
-
批准年份:2013
-
负责人:冯志勇
-
依托单位: