CAREER: Statistical Methods for Dimensionality Reduction in Machine Learning
CAREER: Statistical Methods for Dimensionality Reduction in Machine Learning
批准号:
0650074
负责人:
Lawrence Saul
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-06-30 至 2010-06-30
中文摘要
本研究针对降维问题,发现隐藏在高维数据中的低维结构。 它出现在信息处理的许多领域,并对试图构建模仿人类感知能力的机器的研究人员提出了特别的挑战,例如识别人脸和理解语音。它在统计和科学计算的许多应用中也发挥着越来越重要的作用。随着信息技术的普及,收集和处理越来越多的实验数据成为可能。因此,科学家有兴趣在探索性的分析和可视化的大型多元数据集面临着类似的挑战,在信息处理作为我们的perceptual system.This研究的重点是最近提出的两个算法的维数约简。这两种算法解决了“维数灾难”,因为它出现在两种不同的机器学习环境中:(1)无监督学习,其中在没有来自学习环境的任何反馈的情况下执行降维,以及(2)有监督学习,其中在标记示例的益处下执行降维。要研究的第一个算法是局部线性嵌入(LLE),一种无监督学习算法,用于计算高维数据的低维、邻域保持嵌入。数据,假设躺在一个非线性流形上,被映射到一个单一的全球坐标系的低维。映射是从局部线性重建的对称性导出的,嵌入的实际计算简化为稀疏特征值问题。值得注意的是,LLE中的优化(尽管能够生成高度非线性的嵌入)实现起来很简单,并且它们不涉及局部最小值。LLE在探索性数据分析、科学可视化和计算机视觉中有着广泛的应用。第二种算法是乘法边际最大化(M3),它是支持向量机(SVM)中非负二次规划的监督学习算法。支持向量机目前为机器学习中的许多问题提供了最先进的解决方案,特别是那些涉及高维数据集的问题。然而,解决支持向量机中的二次规划问题仍然是其实现中的一个重要瓶颈。M3算法就是为了缓解这个瓶颈而设计的。它的更新规则具有简单的封闭形式,并且单调收敛于最大间隔超平面的解。此外,它们不涉及任何算法,例如选择学习率或决定在每次迭代时更新哪些变量。它们优化了传统的支持向量机的目标函数,可以应用于分类、回归和新奇检测等问题。本研究所研究的算法易于实现,但它们解决的问题相当复杂。与以前的方法相比,它们的区别不仅在于其新颖的简单性和良好的优化,而且还在于它们与数学,计算机科学和统计学中的其他领域的意想不到的联系。这项工作不仅将发展这些算法的理论基础,还将尝试将其扩展到机器学习中越来越大的问题。这个CAREER奖认可并支持有可能成为二十一世纪学术领袖的教师学者的早期职业发展活动。 该研究预计将通过克服极高维数据集带来的挑战,对科学和工程的许多领域产生广泛影响。软件工具包将被发布,这样各地的研究人员都可以使用最先进的降维方法。教育创新将包括人工智能,机器学习,统计计算和感官处理的新本科和研究生课程。
英文摘要
This research addresses the problem of dimensionality reduction, discovering low dimensional structure hidden in high dimensional data. It arises in many fields of information processing, and poses a particular challenge to researchers attempting to build machines that emulate feats of human perception, such as recognizing faces and understanding speech. It also plays an increasingly prominent role in many applications of statistical and scientific computing. With the advent of widespread information technologies, it has become possible to collect and manipulate ever-increasing amounts of experimental data. Thus, scientists interested in the exploratory analysis and visualization of large multivariate data sets face similar challenges in information processing as our perceptual systems.This research focuses on two recently proposed algorithms for dimensionality reduction. The two algorithms address the "curse of dimensionality" as it arises in two different settings of machine learning: (1) unsupervised learning, where the dimensionality reduction is performed without any feedback from the learning environment, and (2) supervised learning, where the dimensionality reduction is performed with the benefit of labeled examples.The first algorithm to be studied is Locally Linear Embedding (LLE), an unsupervised learning algorithm that computes low dimensional, neighborhood preserving embeddings of high dimensional data. The data, assumed to lie on a nonlinear manifold, is mapped into a single global coordinate system of lower dimensionality. The mapping is derived from the symmetries of locally linear reconstructions, and the actual computation of the embedding reduces to a sparse eigenvalue problem. Notably, the optimizations in LLE (though capable of generating highly nonlinear embeddings) are simple to implement, and they do not involve local minima. LLE has applications to exploratory data analysis, scientific visualization, and computer vision.The second algorithm is Multiplicative Margin Maximization (M3), a supervised learning algorithm for nonnegative quadratic programming in support vector machines (SVMs). Support vector machines currently provide state-of-the-art solutions to many problems in machine learning, particularly those involving data sets of high dimensionality. Solving the quadratic programming problem in SVMs, however, remains a significant bottleneck in their implementation. The M3 algorithm is designed to alleviate this bottleneck. Its update rules have a simple closed form, and they converge monotonically to the solution of the maximum margin hyperplane. Moreover, they do not involve any heuristics such as choosing a learning rate or deciding which variables to update at each iteration. They optimize the traditionally proposed objective function for SVMs and can be applied to problems in classification, regression, and novelty detection.The algorithms to be studied in this research are easy to implement, but the problems they solve are quite complex. Compared to previous approaches, they are distinguished not only by their novel simplicity and well-behaved optimizations, but also by the unexpected connections they make to other areas in mathematics, computer science, and statistics. The work will not only develop the theoretical foundations of these algorithms, but also attempt to scale them up to increasingly large problems in machine learning.This CAREER award recognizes and supports the early career-development activities of a teacher-scholar who is likely to become an academic leader of the twenty-first century. The research is expected to have a broad impact across many areas of science and engineering, by overcoming the challenges posed by data sets of extremely high dimensionality. Software toolkits will be published, so that researchers everywhere will have access to state-of-the-art methods for dimensionality reduction. The educational innovations will include new undergraduate and graduate courses in artificial intelligence, machine learning, statistical computing, and sensory processing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research:EAGER:Deep Architectures for Speech and Audio Processing
-
批准号:0957560
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2010
-
负责人:Lawrence Saul
-
依托单位:
HCC-Small: Assistive Listening Devices and Voice Processing Platforms for the Deaf and Hard of Hearing
-
批准号:0812576
-
项目类别:Continuing Grant
-
资助金额:$14.9万
-
财政年份:2008
-
负责人:Lawrence Saul
-
依托单位:
CAREER: Statistical Methods for Dimensionality Reduction in Machine Learning
-
批准号:0238323
-
项目类别:Continuing grant
-
资助金额:$40.0万
-
财政年份:2003
-
负责人:Lawrence Saul
-
依托单位:
海外基金