课题基金 / 基金详情

CDI Type II: Collaborative Research: Sparse Inference: New Tools for Structural Knowledge Discovery

CDI Type II: Collaborative Research: Sparse Inference: New Tools for Structural Knowledge Discovery
CDI 类型 II:协作研究:稀疏推理:结构知识发现的新工具
批准号:
0835531
负责人:
Bin Yu
金额:
$134.17万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-15 至 2014-02-28

项目摘要

项目成果

Bin Yu的其他基金

相似基金

相关文献

中文摘要
翻译
在人类知识的大多数领域,信息革命导致了大规模的数据爆炸。数据集的规模,其分布和异构性,以及快速提供非专家易于解释的结果的需求,为统计学习提出了新的理论和计算挑战,从变量选择和结构推理到可视化和在线学习。该项目使用稀疏统计推断作为应对这些挑战的强大方法。它的关键见解是,寻求稀疏性是一种同时稳定推理过程和突出底层数据结构的有意义的方法。因此,这项工作将稀疏统计学习的基本进展与数学规划的尖端计算工具相结合,为大规模流数据集中的结构知识发现创建了一个新的框架。它集中在稀疏推理的两个基本主题:变量选择和结构推理。变量选择试图从高维数据集中分离出一些关键变量,是统计学习中的基本预处理工具。然后,结构推理的目的是一致地确定这些变量之间的一些核心依赖关系,以突出其结构。从计算的角度来看,机器学习的许多最新成果都依赖于凸优化的先进方法,如半定规划和鲁棒优化,该项目旨在提高这些算法的复杂性及其处理超大规模流数据的能力。 在实践中,这个项目的动机是希望通过分析大规模的政治和社会数据集,特别是投票记录,在线新闻来源和民意调查数据,帮助公众了解我们的民主。其方法是将统计推断原理应用于社会科学,与政治学和经济学专家合作,建立正在研究的模型和技术。在开展研究时,该项目将把统计学、电气工程/金融工程的研究生培养成统计学、优化和金融学和政治学等学科领域的跨学科研究人员。此外,我们计划开发一个网站,首先访问一组有限的社会科学研究人员,使他们能够分析中等规模的语料库的在线新闻文本格式,在说,稀疏的图表显示给定的关键字之间的统计关联的话。PI计划开发一个软件工具箱来实现这些结果,与MATLAB、R或python等常用数值软件包以及伯克利和普林斯顿大学的“在线数据统计分析”本科课程相连接,将该项目产生的一些材料纳入课程计划。
英文摘要
In most areas of human knowledge, the information revolution has resulted in a massive data explosion. The scale of data sets, their distributed and heterogeneous nature, and the need to quickly deliver results easily interpretable by non-experts, raise new theoretical and computational challenges for statistical learning, from variable selection and structural inference to visualization and online learning. This project uses sparse statistical inference as a powerful approach to meet these challenges. Its key insight is that seeking sparsity is a meaningful way of simultaneously stabilizing inference procedures, and highlighting structure in the underlying data. This work thus combines fundamental advances in sparse statistical learning with cutting-edge computational tools from mathematical programming to create a new framework for structural knowledge discovery in large-scale, streaming data sets. It is focused on two fundamental themes in sparse inference: variable selection and structural inference. Variable selection seeks to isolate a few key variables from high dimensional data sets and is a fundamental preprocessing tool in statistical learning. Structural inference then aims to consistently identify a few core dependence relationships among these variables to highlight its structure. From a computational point of view, many recent results in machine learning have relied on advanced methods from convex optimization such as semidefinite programming and robust optimization and this project seeks to improve the complexity of these algorithms and their capacity to handle very-large scale, streaming data. In practice, this project is motivated by the desire to help the public understand our democracies by analyzing large-scale political and social data sets, with a particular focus on voting records, online news sources, and polling data. Its approach is to apply statistical inference principles to social sciences, using collaborations with experts in political science and economics to forge the models and techniques under study. In carrying out the research, this project will be training graduate students from statistics, electrical engineering/financial engineering into interdisciplinary researchers at the interface of statistics, optimization, and subject matter areas such as finance and political science. In addition we plan to develop a web site, accessible first to a restricted set of social science researchers, to allow them to analyze mid-sized corpora of online news in text format, in the form of say, sparse graphs of words showing statistical associations between given keywords. The PIs plan to develop a software toolbox implementing these results, interfaced with common numerical packages such as MATLAB, R or python as well as an undergraduate course on ``Statistical Analysis of Online Data" at Berkeley and Princeton, incorporating some of the material produced in this project into the course program.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Advancing Theory and Methodology for Tree-Based Algorithms in High Dimensions
  • 批准号:
    2209975
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.0万
  • 财政年份:
    2022
  • 负责人:
    Bin Yu
  • 依托单位:
Understanding Complexity and the Bias-Variance Tradeoff in High Dimensions: Theory and Data Evidence
  • 批准号:
    2015341
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2020
  • 负责人:
    Bin Yu
  • 依托单位:
Parallel Ensemble Learning and Feature Interaction Discovery: High Volume Dynamic Data
  • 批准号:
    1953191
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.2万
  • 财政年份:
    2020
  • 负责人:
    Bin Yu
  • 依托单位:
Understand the functional mechanism of the DSP1 complex in the 3' end maturation of plant small nuclear RNAs
  • 批准号:
    1818082
  • 项目类别:
    Standard Grant
  • 资助金额:
    $68.26万
  • 财政年份:
    2018
  • 负责人:
    Bin Yu
  • 依托单位:
国内基金
海外基金
铋基邻近双金属位点Type B异质结光热催化合成氨机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    30.0万元
  • 批准年份:
    2024
  • 负责人:
    黎景卫
  • 依托单位:
智能型Type-I光敏分子构效设计及其抗耐药性感染研究
  • 批准号:
    22207024
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    20.0万元
  • 批准年份:
    2022
  • 负责人:
    赵琦
  • 依托单位:
TypeⅠR-M系统在碳青霉烯耐药肺炎克雷伯菌流行中的作用机制研究
  • 批准号:
    --
  • 项目类别:
    面上项目
  • 资助金额:
    55万元
  • 批准年份:
    2021
  • 负责人:
    蒋晓飞
  • 依托单位:
替加环素耐药基因 tet(A) type 1 变异体在碳青霉烯耐药肺炎克雷伯菌中的流行、进化和传播
  • 批准号:
    LY22H200001
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2021
  • 负责人:
    蔡加昌
  • 依托单位: