课题基金 / 基金详情

Collaborative Research: MSPA-MCS: Sparse Multivariate Data Analysis

Collaborative Research: MSPA-MCS: Sparse Multivariate Data Analysis
合作研究:MSPA-MCS:稀疏多元数据分析
批准号:
0625409
负责人:
Gert Lanckriet
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-09-15 至 2010-08-31

项目摘要

项目成果

Gert Lanckriet的其他基金

相似基金

相关文献

中文摘要
翻译
本文对经典多变量数据分析方法的稀疏变体进行了发展和研究。它主要关注稀疏主成分分析(PCA)和相关的稀疏典型相关分析(CCA),但也打算探索稀疏变体的方法,如对应分析和判别分析。开发稀疏多元分析算法的动机是它们产生比经典分析更可解释和更健壮的统计结果的潜力,同时在统计效率和表达能力方面尽可能少地放弃。研究人员推导了稀疏PCA的凸松弛作为一个大规模的半确定程序。本研究首先研究了这种松弛的理论和实际性能,以及求解相应半定规划的大规模实例所涉及的计算复杂性。下一步,重点是将这些结果扩展到上面提到的其他多变量数据分析方法。主成分分析(Principal Component Analysis,简称PCA)是一种经典的统计工具,用于研究具有大量变量(气象记录、基因表达系数、利率曲线、社会网络等)的实验数据。它主要用作降维工具:PCA生成一组经过简化的合成变量,这些变量捕获了数据上最大量的信息。这使得在三维图形上表示具有数千个变量的数据集成为可能,同时仍然捕获原始数据的大多数特征,从而使可视化和解释更容易。不幸的是,PCA的主要缺点是这些新的合成变量是所有原始变量的加权和,这使得它们的物理解释很困难。本研究将研究稀疏PCA的计算算法,即在保留原始数据集的大部分特征的同时,计算新的合成变量,这些变量仅是少数问题变量的加权和。稀疏主成分分析是一个很难的组合问题,但研究人员已经产生了一个松弛,可以有效地解决利用最近的结果在凸优化。研究人员计划研究这种松弛的理论和实际性能,并将这些结果扩展到其他统计方法中。
英文摘要
This proposal develops and studies sparse variants of classic multivariate data analysis methods. It primarily focuses on sparse principal component analysis (PCA) and the related sparse canonical correlation analysis (CCA), but also intends to explore sparse variants of methods such as correspondence analysis and discriminant analysis. The motivation for developing sparse multivariate analysis algorithms is their potential for yielding statistical results that are more interpretable and more robust than classical analyses, while giving up as little as possible in the way of statistical efficiency and expressive power. The investigators have derived a convex relaxation for sparse PCA as a large-scale semidefinite program. The proposed research first studies the theoretical and practical performance of this relaxation as well as the computational complexity involved in solving large-scale instances of the corresponding semidefinite programs. In a next step, it focuses on extending these results to the other multivariate data analysis methods cited above. Principal Component Analysis (or PCA) is a classic statistical tool used to study experimental data with a very large number of variables (meteorological records, gene expression coefficients, the interest rate curve, social networks, etc). It is primarily used as a dimensionality reduction tool: PCA produces a reduced set of synthetic variables that captures a maximum amount of information on the data. This makes it possible to represent data sets with thousands of variables on a three dimensional graph while still capturing most of the features of the original data, thus making visualization and interpretation easier.Unfortunately, the key shortcoming of PCA is that these new synthetic variables are a weighted sum of all the original variables making their physical interpretation difficult. The proposed research will study algorithms for computing sparse PCA, i.e., computing new synthetic variables that are the weighted sum of only a few problem variables while keeping most of the features of the original data set. Sparse PCA is a hard combinatorial problem but the investigators have produced a relaxation that can be solved efficiently using recent results in convex optimization. The investigators plan to study the theoretical and practical performance of this relaxation and extend these results to other statistical methods.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: An Integrated Framework for Multimodal Music Search and Discovery
  • 批准号:
    1054960
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $55.0万
  • 财政年份:
    2011
  • 负责人:
    Gert Lanckriet
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)