课题基金 / 基金详情

Robust Estimation and Inference

Robust Estimation and Inference
稳健的估计和推理
批准号:
RGPIN-2014-05227
负责人:
Zamar, Ruben
金额:
$2.04万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Zamar, Ruben的其他基金

相似基金

相关文献

中文摘要
翻译
为了获得有用的推论和预测而必须过滤的误差和扰动来自几个来源,包括:(1)随机波动,例如观测受到测量误差、自然波动和采样变异性的影响;(2)数据污染,例如数据通常包括质量参差不齐的测量、离群值、粗大误差和来自目标群体以外的总体的情况,以及(3)丢失数据。大多数传统的统计程序处理(1),并且有许多论文分别处理(2)和(3)。然而,同时研究(1)、(2)和(3)的论文很少。我提出的一些研究将旨在填补这一空白。我希望开发能够使用计算效率和可伸缩算法来处理上述所有不确定性来源的程序。*考虑一个数据表,它有n行--每种情况一行--和p列--每一变量或特征一列。随着廉价计算和存储的出现,许多现代数据集变量丰富,案例匮乏。这在文献中被称为“小n-大p问题”。这一现象也与统计学中所谓的维度灾难问题有关。给定某个目标(例如,预测数据表中某个响应变量(S)的未来值),通常会发现大量变量(我称之为噪声变量)对此任务没有帮助,反而有害。因此,噪声变量构成了第四类扰动,需要对其进行滤波,以更好地提取剩余信号变量中包含的信息。此外,信号变量本身可能是部分冗余的,并且信号变量的子集(我们称为方阵)可能比信号变量的全集具有更好的预测能力。方阵可以用来构建统计模型,然后可以对结果进行集成,以提供单一的预测/分类。选择方阵(方阵编队)的问题是模型选择的泛化,我们允许不同的变量组形成协作模型来执行单个任务。关于这种建立模式的方法,有许多实际和理论问题,我想谈一谈。我们以前的博士生Jabed Tomal在药物发现的背景下,在这个主题上做了一些开创性的工作。韦尔奇教授和我现在希望招收一名新的博士生来扩展这项工作,这项工作有可能在工业和科学的许多领域应用。*经典的稳健性模型基于这样的范例,即绝大多数情况(数据表中的行)没有污染,对执行给定的任务是有用的。因此,可能只需要确定和筛选(降低权重)少数受污染病例。遗憾的是,对于非常高维的数据表,这种范例并不完全令人满意。如果单元格(数据表中的单个条目)受到污染的概率d很小,则案例(数据表中的一行)受到污染的概率为e=1-(1-d)^{p},它可能很快变得大于0.5。例如,如果d=0.63397,p=100,则e=100。Alqallaf,Van Aelst,Yohai和Zamar(2009)引起了人们对“异常值传播”这个问题的关注,并提出了一些可能的方法来解决这个问题。我希望进一步研究这个问题。我以前的博士生Mike Danilov构建了稳健的S-多变量位置和散布的估计,可以有效地处理随机单元格中的缺失。这是构建针对异常值传播的稳健估计的重要构件。我目前的博士生梁安迪正在追寻这个研究方向。
英文摘要
Errors and perturbations which must be filtered to obtain useful inferences and predictions arise from several sources, including: (1) random fluctuations, e.g. observations are affected by measurement errors, natural fluctuations and sampling variability, (2) data contamination, e.g. data often include measurements of uneven quality, outliers, gross errors and cases from populations other than the target one, and (3) missing data. Most traditional statistical procedures deal with (1) and there are many papers dealing with (2) and (3) separately. However, there are few papers dealing with(1), (2) and (3) simultaneously. Some of my proposed research will aim at filling this gap. I wish to develop procedures able to deal with all the above mentioned sources of uncertainty, using computational efficient and scalable algorithms.* Consider a data table with n rows -- one for each case -- and p columns -- one for each variable or feature. With the advent of cheap computing and storage, many modern datasets are variables-rich and cases-poor. This is referred to as "small n-- large p problem" in the literature. This phenomenon is also related to the so called curse of dimensionality problem in Statistics. Given a certain goal (e.g. prediction of future values for some response variable (s) in the data table, it is common to find that a large number of variables (which I call noise variables) hurt instead of helping this task. Hence noise variables constitute a fourth type of perturbation which needs to be filtered to better extract the information contained in the remaining signal variables. In addition, signal variables themselves may be partially redundant and subsets of signal variables (which we call phalanxes) may have better predictive power than the full set of signal variables. Phalanxes can be used to construct statistical models which results can then be ensembled to provide a single prediction/classification. The problem of selecting phalanxes (phalanx formation) is a generalization of model selection where we allow for different groups of variables to form cooperating models to perform a single task. There are many practical and theoretical questions regarding this model building approach which I would like to address. Our former PhD student Jabed Tomal did some ground breaking work on this topic in the context of drug discovery. Prof. Welch and I now wish to enroll a new PhD student to expand this work which has potential for application in many areas of industry and science.* The classical robustness model is based on the paradigm that the vast majority of cases (rows in the data table) are free of contamination and useful to perform the given task. Hence, only a minority of contaminated cases may need to be identified and filtered (downweighted). Unfortunately this paradigm is not fully satisfactory in the case of very high dimensional data tables. If there is a small and independent probability, d, that a cell (individual entry in the data table) is contaminated then the probability that a case (a row in the data table) is contaminated is e=1- (1-d)^{p} which can quickly become larger than 0.5. For example, if d=0.01 and p=100 we have e=0.63397. Alqallaf, Van Aelst, Yohai and Zamar (2009) brings attention to this problem called "propagation of outliers" and propose some possible approaches to address it. I wish to further study this problem. My former Ph.D. student Mike Danilov constructed robust S-estimates of multivariate location and scatter that can efficiently deal with missing at random cells. This was an important building block for constructing robust estimates against outliers propagation. My current PhD student Andy Leung is pursuing this research direction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Robust Estimation and Model Ensemble Selection
  • 批准号:
    RGPIN-2019-04201
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.82万
  • 财政年份:
    2022
  • 负责人:
    Zamar, Ruben
  • 依托单位:
Robust Estimation and Model Ensemble Selection
  • 批准号:
    RGPIN-2019-04201
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.82万
  • 财政年份:
    2021
  • 负责人:
    Zamar, Ruben
  • 依托单位:
Robust Estimation and Model Ensemble Selection
  • 批准号:
    RGPIN-2019-04201
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.82万
  • 财政年份:
    2020
  • 负责人:
    Zamar, Ruben
  • 依托单位:
Robust Estimation and Model Ensemble Selection
  • 批准号:
    RGPIN-2019-04201
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.82万
  • 财政年份:
    2019
  • 负责人:
    Zamar, Ruben
  • 依托单位:
海外基金