CDI-Type II: Collaborative Research: Sparse Inference: New Tools for Structural Knowledge Discovery
CDI-Type II: Collaborative Research: Sparse Inference: New Tools for Structural Knowledge Discovery
批准号:
0835550
负责人:
Laurent El Ghaoui
金额:
$41.72万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-15 至 2013-02-28
中文摘要
在人类知识的大多数领域,信息革命导致了大量数据爆炸。数据集的规模,它们的分布和异构性质,以及快速提供非专家易于解释的结果的需求,为统计学习提出了新的理论和计算挑战,从变量选择和结构推断到可视化和在线学习。本项目使用稀疏统计推断作为应对这些挑战的强大方法。它的关键观点是,寻求稀疏性是同时稳定推理过程和突出底层数据结构的一种有意义的方法。因此,这项工作将稀疏统计学习的基本进展与来自数学规划的尖端计算工具相结合,为大规模流数据集中的结构知识发现创建了一个新的框架。它集中在稀疏推理的两个基本主题:变量选择和结构推理。变量选择旨在从高维数据集中分离出几个关键变量,是统计学习中的基本预处理工具。然后,结构推理旨在一致地识别这些变量之间的几个核心依赖关系,以突出其结构。从计算的角度来看,机器学习的许多最新成果都依赖于来自凸优化的先进方法,如半定规划和鲁棒优化,该项目旨在提高这些算法的复杂性及其处理大规模流数据的能力。在实践中,这个项目的动机是通过分析大规模的政治和社会数据集,特别关注投票记录、在线新闻来源和投票数据,帮助公众了解我们的民主制度。它的方法是将统计推理原理应用于社会科学,与政治学和经济学专家合作,形成所研究的模型和技术。在开展这项研究时,该项目将把统计学、电气工程/金融工程的研究生培养成统计学、优化以及金融和政治学等学科领域的跨学科研究人员。此外,我们计划开发一个网站,首先供有限的社会科学研究人员访问,允许他们以文本格式分析中等规模的在线新闻语料库,例如,以显示给定关键词之间统计关联的稀疏图形的形式。pi计划开发一个实现这些结果的软件工具箱,与MATLAB、R或python等通用数值软件包以及伯克利和普林斯顿大学的“在线数据统计分析”本科课程相结合,将该项目中产生的一些材料纳入课程计划。
英文摘要
In most areas of human knowledge, the information revolution has resulted in a massive data explosion. The scale of data sets, their distributed and heterogeneous nature, and the need to quickly deliver results easily interpretable by non-experts, raise new theoretical and computational challenges for statistical learning, from variable selection and structural inference to visualization and online learning. This project uses sparse statistical inference as a powerful approach to meet these challenges. Its key insight is that seeking sparsity is a meaningful way of simultaneously stabilizing inference procedures, and highlighting structure in the underlying data. This work thus combines fundamental advances in sparse statistical learning with cutting-edge computational tools from mathematical programming to create a new framework for structural knowledge discovery in large-scale, streaming data sets. It is focused on two fundamental themes in sparse inference: variable selection and structural inference. Variable selection seeks to isolate a few key variables from high dimensional data sets and is a fundamental preprocessing tool in statistical learning. Structural inference then aims to consistently identify a few core dependence relationships among these variables to highlight its structure. From a computational point of view, many recent results in machine learning have relied on advanced methods from convex optimization such as semidefinite programming and robust optimization and this project seeks to improve the complexity of these algorithms and their capacity to handle very-large scale, streaming data. In practice, this project is motivated by the desire to help the public understand our democracies by analyzing large-scale political and social data sets, with a particular focus on voting records, online news sources, and polling data. Its approach is to apply statistical inference principles to social sciences, using collaborations with experts in political science and economics to forge the models and techniques under study. In carrying out the research, this project will be training graduate students from statistics, electrical engineering/financial engineering into interdisciplinary researchers at the interface of statistics, optimization, and subject matter areas such as finance and political science. In addition we plan to develop a web site, accessible first to a restricted set of social science researchers, to allow them to analyze mid-sized corpora of online news in text format, in the form of say, sparse graphs of words showing statistical associations between given keywords. The PIs plan to develop a software toolbox implementing these results, interfaced with common numerical packages such as MATLAB, R or python as well as an undergraduate course on ``Statistical Analysis of Online Data" at Berkeley and Princeton, incorporating some of the material produced in this project into the course program.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Mathematical Programming for Streaming Data
-
批准号:1250687
-
项目类别:Standard Grant
-
资助金额:$15.05万
-
财政年份:2011
-
负责人:Laurent El Ghaoui
-
依托单位:
Collaborative Research: Mathematical Programming for Streaming Data
-
批准号:0969923
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2010
-
负责人:Laurent El Ghaoui
-
依托单位:
Collaborative Research: Mathematical Programming for Streaming Data
-
批准号:0968842
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2010
-
负责人:Laurent El Ghaoui
-
依托单位:
Collaborative Research: MSPA-MCS: Sparse Multivariate Data Analysis
-
批准号:0625371
-
项目类别:Standard Grant
-
资助金额:$23.0万
-
财政年份:2006
-
负责人:Laurent El Ghaoui
-
依托单位:
CAREER: Robust Optimization and Applications
-
批准号:9983874
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2000
-
负责人:Laurent El Ghaoui
-
依托单位:
国内基金
海外基金
登录
查看更多内容
铋基邻近双金属位点Type B异质结光热催化合成氨机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:30.0万元
-
批准年份:2024
-
负责人:黎景卫
-
依托单位:
智能型Type-I光敏分子构效设计及其抗耐药性感染研究
-
批准号:22207024
-
项目类别:青年科学基金项目(C类)
-
资助金额:20.0万元
-
批准年份:2022
-
负责人:赵琦
-
依托单位:
TypeⅠR-M系统在碳青霉烯耐药肺炎克雷伯菌流行中的作用机制研究
-
批准号:--
-
项目类别:面上项目
-
资助金额:55万元
-
批准年份:2021
-
负责人:蒋晓飞
-
依托单位:
替加环素耐药基因 tet(A) type 1 变异体在碳青霉烯耐药肺炎克雷伯菌中的流行、进化和传播
-
批准号:LY22H200001
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:蔡加昌
-
依托单位:
面向手性α-氨基酰胺药物的新型不对称Ugi-type 反应开发
-
批准号:LY22B020003
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:李绍玉
-
依托单位:
BMP9/BMP type I receptors 通过激活 PPARα保护心肌梗死的机制研究
-
批准号:LQ22H020003
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:陈灵丽
-
依托单位:
C2H2-type锌指蛋白在香菇采后组织软化进程中的作用机制研究
-
批准号:32102053
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:邓冰
-
依托单位:
血管阻断型Type-I光敏剂合成及其三阴性乳腺癌光诊疗
-
批准号:62120106002
-
项目类别:国际(地区)合作与交流项目
-
资助金额:255万元
-
批准年份:2021
-
负责人:董晓臣
-
依托单位:
茶尺蠖Type-II环氧性信息素合成酶关键基因的鉴定及功能研究
-
批准号:LQ21C140001
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2020
-
负责人:王倩
-
依托单位:
Chichibabin-type偶联反应在构建联氮杂芳烃中的应用
-
批准号:22078300
-
项目类别:面上项目
-
资助金额:63.0万元
-
批准年份:2020
-
负责人:李景华
-
依托单位: