课题基金 / 基金详情

Beyond Clustering: Unsupervised Modeling with Complex Representations

Beyond Clustering: Unsupervised Modeling with Complex Representations
超越聚类:具有复杂表示的无监督建模
批准号:
EP/E042694/1
负责人:
Katherine Heller
金额:
$29.72万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2008
资助国家:
英国
项目状态:
已结题
起止时间:
2008 至 --

项目摘要

项目成果

Katherine Heller的其他基金

相似基金

相关文献

中文摘要
翻译
机器学习领域致力于开发新的理论和算法,以提高计算机识别模式、做出自主决策和基于数据进行预测的能力。机器学习的新进展在其他科学领域、商业和个人日常生活中都产生了广泛的影响。例如,它们可以帮助神经科学家分析高维大脑成像数据,改进在线产品推荐系统,或者帮助个人自动组织他们的数字相册。聚类是解决各种问题的重要的无监督机器学习工具。抽象地说,集群是发现属于一起的数据点的组。举个例子,如果给动物分组的任务,人们可能会按类型(哺乳动物、爬行动物、两栖动物)或按大小(小或大)将它们分组。自动聚类工具已被用于对基因表达数据进行聚类以阐明基因功能,自动按主题对网络新闻文章进行分组,自动按流派对音乐进行分类,以及对气候数据进行时空聚类以改进气候预测。尽管聚类对于许多应用来说是一个很好的工具,但实际上它是相当有限的。在许多情况下,与简单地将每个数据点分配给集群相比,被建模的数据可以具有更丰富、更复杂的隐藏表示。例如,数据点实际上可以同时属于多个集群(例如,电影《呐喊》可以同时属于恐怖电影集群和喜剧集群)。数据的隐藏表示可以是结构化的,例如,句子可以由句法分析树表示。正在建模的数据可能具有多个潜在特征(如可以包含多个对象的图像)。此外,潜在特征的总数可能是未知的,因此不应预先指定或限制。这种灵活性是通过使用非参数贝叶斯方法提供的,这将在本提议中发挥基础作用。我的主要目标是通过开发原则性的、理论上合理的概率模型和算法来促进无监督机器学习的最新发展,这些模型和算法将聚类范例扩展到需要更丰富表示的问题。这些更丰富、更复杂的数据表示提供了在集群不够好的许多情况下很好地对数据建模的能力。除了推进理论,我还将为使用这些表示法的概率模型开发高效的学习和推理算法。我大部分工作的起点将是非参数贝叶斯方法,特别是印度巴菲特过程(IBP)。非参数方法被设计为非常灵活,可以比固定参数数量的不灵活模型更好地建模数据。我的方法将能够从数据中自动推断出正确的模型大小(参数数量)。我将重点介绍无监督机器学习的六个具体新贡献。首先,我将开发概率模型,其中每个数据点可以同时属于多个重叠的集群。其次,我将把按需聚类范例扩展到关系数据,创建一种方法,使计算机能够执行简单形式的类比推理。第三,我将开发有效的IBPS学习和推理方法。第四,使用IBP I将创建一种新的独立分量分析方法(一种广泛使用的信号处理方法),使其能够自动学习信号中的分量数量。第五,我将为计算机开发新的概率非监督方法,将他们在一项任务中学到的知识转移到其他任务中。最后,我将探索高级概率理论和随机过程在设计实用的非参数机器学习方法中的新用途。
英文摘要
The field of Machine Learning strives to develop new theory and algorithms that improve the ability of computers to recognize patterns, make autonomous decisions, and make predictions based on data. New advances in Machine Learning have broad impact in other scientific fields, in commerce, and in the daily lives of individuals. For example, they can help neuroscientists analyze high-dimensional brain imaging data, improve online product recommendation systems, or help individuals automatically organize their digital photo albums.Clustering is an important unsupervised Machine Learning tool for a variety of problems. Abstractly, clustering is discovering groups of data points that belong together. As an example, if given the task of clustering animals, one might group them together by type (mammals, reptiles, amphibians), or alternatively by size (small or large). Automated clustering tools have been used to cluster gene expression data in order to elucidate gene function, automatically group news articles on the web by topic, automatically categorize music by genre, and spatio-temporally cluster climate data to improve climate prediction.While clustering is a wonderful tool for many applications, it is actually quite limited. In many situations the data being modeled can have a much richer and more complex hidden representation than the simple assignment of each data point to a cluster. For example, data points can actually belong to multiple clusters simultaneously (e.g. the movie Scream could belong to both the horror movie cluster and the comedy cluster). The hidden representation of the data could be structured, for example sentences can be represented by parse trees. The data being modeled might have multiple latent features (like images which can contain multiple objects). Moreover, the total number of latent features might not be known, and therefore should not be specified or limited a priori. This flexibility is provided by the use of nonparametric Bayesian methods, which will play a fundamental role in this proposal.My main goal is to advance the state-of-the-art for unsupervised machine learning, by developing principled, theoretically sound, probabilistic models and algorithms, which extend a clustering paradigm to problems which need richer representations. These richer and more complex representations for data provide the ability to model data well in the many situations in which clustering is not good enough. In addition to advancing the theory, I will also develop efficient learning and inference algorithms for the probabilistic models that use these representations.The starting point for much of my work will be nonparametric Bayesian methods, and in particular, the Indian Buffet Process (IBP). Nonparametric methods are designed to be very flexible, and can model data better than inflexible models with a fixed number of parameters. My methods will be able to automatically infer the correct model size (number of parameters) from the data. I will focus on six specific new contributions to unsupervised machine learning. First, I will develop probabilistic models in which each data point can simultaneously belong to multiple overlapping clusters. Second, I will extend the clustering-on-demand paradigm to relational data creating a method that will enable computers to perform simple forms of analogical reasoning. Third, I will develop efficient methods for learning and inference in IBPs. Fourth, using the IBP I will create a new approach to Independent Components Analysis (a widely-used signal processing method) making it possible to automatically learn the number of components in a signal. Fifth, I will develop new probabilistic unsupervised methods for computers to transfer what they have learned on one task to other tasks. Finally, I will explore new uses of advanced probability theory and stochastic processes in the design of practical nonparametric machine learning methods.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Interacting Dynamic Bayesian Models for Social Behavior and Reasoning
  • 批准号:
    1553465
  • 项目类别:
    Standard Grant
  • 资助金额:
    $51.6万
  • 财政年份:
    2016
  • 负责人:
    Katherine Heller
  • 依托单位:
BRAIN EAGER: Integrative Cross-Modal and Cross-Species Brain Models: Motivation and Reward
  • 批准号:
    1451017
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2014
  • 负责人:
    Katherine Heller
  • 依托单位:
Bayesian Models of Social Behavior Using Online Resources
  • 批准号:
    1339593
  • 项目类别:
    Standard Grant
  • 资助金额:
    $7.85万
  • 财政年份:
    2013
  • 负责人:
    Katherine Heller
  • 依托单位:
Workshop for Women in Machine Learning
  • 批准号:
    1346800
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.0万
  • 财政年份:
    2013
  • 负责人:
    Katherine Heller
  • 依托单位:
海外基金