Learning Better Representations With Kernels
Learning Better Representations With Kernels
批准号:
RGPIN-2021-02974
负责人:
Sutherland, Danica
金额:
$2.11万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Many of the most impressive gains in machine learning in the past decade - helping computers "understand" images, text, and more - have been based on the idea of representation learning. Rather than processing an image as a set of red, green, and blue values for a grid of pixels, we learn a set of numbers describing the image in some abstract way, so that two similar images - which may have very different pixel values - have similar representations. Most work on representation learning, however, assumes that we will use the representation only in a very simple (linear) way. But in some settings, especially when instead of trying to find a representation for one static problem we want to find a representation that works for many related problems, this assumption can make our job much harder. A machine learning approach known as kernel methods provides richer ways to work with a representation. We can, then, potentially use simpler, easier-to-find representations than if we force ourselves to use them linearly - especially in settings where we want to find representations good for more than one thing. This scheme has seen good success in several areas already, including in generative models (training a computer to output, e.g., images of fake people) and in density estimation (figuring out the "shape" of a dataset, to know which kinds of points we might expect to see in the future). We believe that it can be applied more widely, to improve our ability to find representations in a variety of problems. For instance, a method known as "invariant risk minimization" tries to find predictors which don't rely on random correlations that happen to be present in the training data, but may not hold when we go to apply the model on slightly different data. This method currently works with only linear predictions based on a given representation, and that assumption can cause it to behave very poorly even on some extremely simple datasets. We believe that a version of the method that incorporates kernel models will be more robust and reliable in applications. One area in particular that has already benefited from this approach is called two-sample testing: telling whether two different datasets are fundamentally different from one another, not just different due to random chance. For example, this is used to tell whether the control group and treatment group are different in a medical trial. In practice, though, it's often the case that the two datasets are different - but only in ways we don't really care about. The methods to help understand how two datasets differ, rather than just whether they're different, are much less developed. We believe that we can use kernels to find representations that will, in the end, help people actually understand the differences between datasets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金