Beyond Clustering: Unsupervised Modeling with Complex Representations
Beyond Clustering: Unsupervised Modeling with Complex Representations
批准号:
EP/E042694/1
负责人:
Katherine Heller
金额:
$29.72万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2008
资助国家:
英国
项目状态:
已结题
起止时间:
2008 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The field of Machine Learning strives to develop new theory and algorithms that improve the ability of computers to recognize patterns, make autonomous decisions, and make predictions based on data. New advances in Machine Learning have broad impact in other scientific fields, in commerce, and in the daily lives of individuals. For example, they can help neuroscientists analyze high-dimensional brain imaging data, improve online product recommendation systems, or help individuals automatically organize their digital photo albums.Clustering is an important unsupervised Machine Learning tool for a variety of problems. Abstractly, clustering is discovering groups of data points that belong together. As an example, if given the task of clustering animals, one might group them together by type (mammals, reptiles, amphibians), or alternatively by size (small or large). Automated clustering tools have been used to cluster gene expression data in order to elucidate gene function, automatically group news articles on the web by topic, automatically categorize music by genre, and spatio-temporally cluster climate data to improve climate prediction.While clustering is a wonderful tool for many applications, it is actually quite limited. In many situations the data being modeled can have a much richer and more complex hidden representation than the simple assignment of each data point to a cluster. For example, data points can actually belong to multiple clusters simultaneously (e.g. the movie Scream could belong to both the horror movie cluster and the comedy cluster). The hidden representation of the data could be structured, for example sentences can be represented by parse trees. The data being modeled might have multiple latent features (like images which can contain multiple objects). Moreover, the total number of latent features might not be known, and therefore should not be specified or limited a priori. This flexibility is provided by the use of nonparametric Bayesian methods, which will play a fundamental role in this proposal.My main goal is to advance the state-of-the-art for unsupervised machine learning, by developing principled, theoretically sound, probabilistic models and algorithms, which extend a clustering paradigm to problems which need richer representations. These richer and more complex representations for data provide the ability to model data well in the many situations in which clustering is not good enough. In addition to advancing the theory, I will also develop efficient learning and inference algorithms for the probabilistic models that use these representations.The starting point for much of my work will be nonparametric Bayesian methods, and in particular, the Indian Buffet Process (IBP). Nonparametric methods are designed to be very flexible, and can model data better than inflexible models with a fixed number of parameters. My methods will be able to automatically infer the correct model size (number of parameters) from the data. I will focus on six specific new contributions to unsupervised machine learning. First, I will develop probabilistic models in which each data point can simultaneously belong to multiple overlapping clusters. Second, I will extend the clustering-on-demand paradigm to relational data creating a method that will enable computers to perform simple forms of analogical reasoning. Third, I will develop efficient methods for learning and inference in IBPs. Fourth, using the IBP I will create a new approach to Independent Components Analysis (a widely-used signal processing method) making it possible to automatically learn the number of components in a signal. Fifth, I will develop new probabilistic unsupervised methods for computers to transfer what they have learned on one task to other tasks. Finally, I will explore new uses of advanced probability theory and stochastic processes in the design of practical nonparametric machine learning methods.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Interacting Dynamic Bayesian Models for Social Behavior and Reasoning
-
批准号:1553465
-
项目类别:Standard Grant
-
资助金额:$51.6万
-
财政年份:2016
-
负责人:Katherine Heller
-
依托单位:
BRAIN EAGER: Integrative Cross-Modal and Cross-Species Brain Models: Motivation and Reward
-
批准号:1451017
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2014
-
负责人:Katherine Heller
-
依托单位:
Bayesian Models of Social Behavior Using Online Resources
-
批准号:1339593
-
项目类别:Standard Grant
-
资助金额:$7.85万
-
财政年份:2013
-
负责人:Katherine Heller
-
依托单位:
Workshop for Women in Machine Learning
-
批准号:1346800
-
项目类别:Standard Grant
-
资助金额:$4.0万
-
财政年份:2013
-
负责人:Katherine Heller
-
依托单位:
Bayesian Models of Social Behavior using Online Resources
-
批准号:1048563
-
项目类别:Standard Grant
-
资助金额:$24.0万
-
财政年份:2011
-
负责人:Katherine Heller
-
依托单位:
海外基金