课题基金 / 基金详情

CAREER: Nonparametric Models Building, Estimation, and Selection with Applications to High Dimensional Data Mining

CAREER: Nonparametric Models Building, Estimation, and Selection with Applications to High Dimensional Data Mining
职业:非参数模型构建、估计和选择及其在高维数据挖掘中的应用
批准号:
0645293
负责人:
Hao Zhang
金额:
$40.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-07-01 至 2013-09-30

项目摘要

项目成果

Hao Zhang的其他基金

相似基金

相关文献

中文摘要
翻译
非参数方法越来越多地应用于回归、分类和密度估计,无论是在统计学领域还是在其他相关领域,如数据挖掘和机器学习。然而,由于维度灾难的存在,非参数模型的一个关键困难是对高维数据的模型拟合。另一个困难是模型推理和解释,即如何评估或测试复杂曲面拟合上的个体变量效应。对于具有复杂协方差结构的异质数据,非参数模型估计更具挑战性。该建议的目的是开发新的和广泛适用的程序来同时选择和估计数据挖掘中非参数模型及其相关范例的模型。在再生核Hilbert空间(RKHS)的框架下,PI针对几类模型提出了一系列新的正则化技术:相关数据的平滑样条方差模型,半参数回归模型,监督和半监督学习的支持向量机。所提出的方法是相对于标准方法的关键进步,因为它们在实现模型稀疏性和函数光滑化方面具有统一的框架,它们的理论属性易于处理,并且它们很容易适应高维问题。PI将研究建议的估计器的渐近行为,探索数据驱动的调整正则化参数的过程,并开发计算算法和软件来实施建议的过程。PI还将通过广泛的模拟研究和真实数据分析来检验新方法的有限样本性能。在当前的信息时代,科学和工业数据库的数量和复杂性已经呈指数级增长。因此,数据表单的维度越来越高。这些数据的分析对统计学家提出了新的挑战,并正在成为现代统计学中最重要的研究课题之一。该项目的目的是显著增加可用于分析复杂的高维数据的工具。在这个项目中,PI旨在实现以下三个目标:(1)在统一的数学框架内应对非参数模型估计和选择的挑战;(2)开发具有所需统计特性的灵活方法和用于挖掘海量数据的高性能统计软件;(3)将上述两个活动的研究机会和成果整合到研究生、本科和高中水平的学科和跨学科统计教育中。这项研究将拓宽对非参数推断和模型选择的传统理解,为包括社会学、经济学、环境、生物和医学在内的各个领域的广泛研究人员和从业者提供最先进的数据分析工具,并帮助下一代学生准备必要的现代统计学观点。
英文摘要
Nonparametric methods are increasingly applied to regression, classification and density estimation, both in statistics and other related areas such as data mining and machine learning. However, a key difficulty with nonparametric models is model fitting for high dimensional data due to the curse of dimensionality. Another difficulty is model inference and interpretation, i.e., how to evaluate or test individual variable effects on the complex surface fit. For heterogeneous data with complicated covariance structure, nonparametric model estimation is even more challenging. The objectives of this proposal are to develop novel and widely applicable procedures to simultaneous model selection and estimation for nonparametric models and their related paradigms in data mining. In the framework of reproducing kernel Hilbert space (RKHS), the PI proposes a host of new regularization techniques for several families of models: smoothing spline ANOVA models for correlated data, semiparametric regression models, support vector machines for supervised and semi-supervised learning. The proposed methodologies constitute key advances over standard methods through their unified framework for achieving model sparsity and function smoothing altogether, their tractable theoretical properties, and their easy adaptation to high dimensional problems. The PI will study asymptotic behaviors of the proposed estimators, explore data-driven procedures for tuning regularization parameters, and develop computation algorithms and softwares to implement the proposed procedures. The PI will also examine finite sample performance of new methods via extensive simulation studies and real data analysis.In the current information era, the volume and complexity of scientific and industrial databases have been exponentially expanding. As a consequence, the data form keeps gaining higher and higher dimensionality. Analysis of such data poses new challenges to statisticians and is becoming one of the most important research topics in modern statistics. The purpose of this project is to significantly increase the available tools for analyzing complex high dimensional data. In this project, the PI aims to accomplish the following three goals: (1) meet the challenges of nonparametric model estimation and selection within a unified mathematical framework; (2) develop flexible methods with desired statistical properties and high-performance statistical softwares for mining massive data; (3) integrate research opportunities and findings from the above two activities into disciplinary and interdisciplinary statistical education at graduate, undergraduate and high school levels. This research will broaden traditional understanding of nonparametric inferences and model selection, provide a broad range of researchers and practitioners in various fields including sociology, economics, environmental, biological and medical sciences with state-of-the-art data analysis tools, and help to prepare the next-generation students with the necessary modern statistical perspectives.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Robot Reflection in Lifelong Adaptation
  • 批准号:
    2308492
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2022
  • 负责人:
    Hao Zhang
  • 依托单位:
CAREER: Robot Reflection in Lifelong Adaptation
  • 批准号:
    1942056
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2020
  • 负责人:
    Hao Zhang
  • 依托单位:
Spectroscopic photon localization microscopy for super-resolution molecular imaging
  • 批准号:
    1706642
  • 项目类别:
    Standard Grant
  • 资助金额:
    $58.62万
  • 财政年份:
    2017
  • 负责人:
    Hao Zhang
  • 依托单位:
TRIPODS: UA-TRIPODS - Building Theoretical Foundations for Data Sciences
  • 批准号:
    1740858
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $136.85万
  • 财政年份:
    2017
  • 负责人:
    Hao Zhang
  • 依托单位:
海外基金