课题基金 / 基金详情

High-dimensional Data Analysis: Modeling Unobserved Heterogeneity in Data, and Studying Imbalanced Classification Problems

High-dimensional Data Analysis: Modeling Unobserved Heterogeneity in Data, and Studying Imbalanced Classification Problems
高维数据分析:对数据中未观察到的异质性进行建模,并研究不平衡分类问题
批准号:
RGPIN-2020-05011
负责人:
Khalili, Abbas
金额:
$1.75万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Khalili, Abbas的其他基金

相似基金

相关文献

中文摘要
翻译
由于当今世界数据收集手段的不断扩大,数据科学已成为广泛科学学科的关注中心。在许多应用中,当前数据前所未有的大小和结构复杂性要求从这些数据中提取有用信息的计算效率和统计上合理的方法。为了实现这一目标,我的研究计划的总主题是分析高维数据。更具体地说,在这项提议的五年中,我的短期目标是:i)异质高维数据的统计建模:在健康科学、工程与环境、社会科学和金融计量经济学等应用中,高维数据往往来自由多个隐藏的同质子总体组成的异质总体。有限混合回归(FMR)和马尔可夫状态切换自回归(MSAR)模型为捕捉数据中未观察到的异质性提供了灵活的工具。后一种模型用于时间序列数据建模。在实践中,当将这样的模型拟合到数据集时,人们面临三个推断问题:顺序选择或隐藏子总体或制度的数量估计、变量选择、以及所谓的选择后统计推断,例如对数据驱动的选定模型的参数的假设检验或可信区间。尽管它们得到了广泛的应用,但在关于高维统计的日益增长的文献中,针对上述问题的严格的方法学发展一直非常有限。在我的短期目标中,我将研究新的基于似然的正则化技术,用于:FMR和MSAR中的阶数选择,以及固定阶数和高维环境下稀疏动态FMR和矢量MSAR中的变量选择。建立这样的结果将为我长期目标的主题--选拔后推理问题铺平道路。Ii)高维不平衡分类问题:在欺诈检测、医疗诊断或设备故障检测等应用中,分类任务经常受到高维和训练数据中某些类别观察频率不平衡的影响。后者要么是由于数据收集过程,要么是因为某些班级在人口中确实很少见。由于少数类数据的稀缺性,传统的判别方法往往偏向多数类,导致少数类的误识率较高。不平衡分类问题通常很难,所以我从研究不平衡线性二进制情况开始。我将研究在针对高维少数民族的标准线性判别分析中,将分而治之的技术与硬阈值变量选择方法相结合进行偏差校正的实用性。我还会研究多班级问题。
英文摘要
Data science has become the center of attention in a wide range of scientific disciplines, thanks to ever-expanding means of data collection in today's world. Unprecedented size and structural complexity of current data in many applications call for computationally efficient and statistically sound methodologies for extracting useful information from such data. Toward this goal, the general theme of my research program focuses on analyzing high-dimensional data. More specifically, over the five years of this proposal, my short-term objectives are: I) Statistical modeling of heterogeneous high-dimensional data: In applications such as health sciences, engineering and environment, social sciences, and financial econometrics, high-dimensional data often arise from heterogeneous populations consisting of multiple hidden homogeneous sub-populations. Finite mixture of regressions (FMR) and Markov regime-switching autoregressive (MSAR) models provide flexible tools for capturing unobserved heterogeneity in data. The later models are used for modeling time series data. In practice, when fitting such models to a dataset, one faces three inferential problems: order selection or estimation of the number of hidden sub-populations or regimes, variable selection, and so-called post-selection statistical inference such as hypothesis testing or confidence intervals for parameters of a data-driven selected model. Despite their wide applications, rigorous methodological developments addressing the aforementioned problems in the growing literature on high-dimensional statistics have been very limited. In my short-term objectives, I will investigate new likelihood-based regularization techniques for: order selection in FMR and MSAR, and variable selection in sparse dynamic FMR and vector MSAR with fixed order and in high-dimensional settings. Establishment of such results will pave the way toward post-selection inference problems which are the subjects of my long-term objectives. II) High-dimensional imbalanced classification problems: In applications such as fraud detection, medical diagnosis, or equipment malfunction detection, classification tasks often suffer from both high-dimensionality and imbalance in the observed frequency of some classes in the training data. The latter is due to either data collection process or because some classes are indeed rare in the population. Due to data scarcity in minority class(es), conventional discriminative methods are often biased toward the majority class(es) resulting in much higher misclassification rates for the minority class(es). Imbalanced classification problems are generally hard, so I begin by studying imbalanced linear binary cases. I will investigate the utility of divide-and-conquer techniques coupled with hard-thresholding variable selection methods for bias correction in the standard linear discriminant analysis toward the minority class in high-dimensions. I will also study multi-class problems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
High-dimensional Data Analysis: Modeling Unobserved Heterogeneity in Data, and Studying Imbalanced Classification Problems
  • 批准号:
    RGPIN-2020-05011
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.75万
  • 财政年份:
    2021
  • 负责人:
    Khalili, Abbas
  • 依托单位:
High-dimensional Data Analysis: Modeling Unobserved Heterogeneity in Data, and Studying Imbalanced Classification Problems
  • 批准号:
    RGPIN-2020-05011
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.75万
  • 财政年份:
    2020
  • 负责人:
    Khalili, Abbas
  • 依托单位:
Statistical inference in finite mixture of regressions and mixture-of-experts models in high-dimensional spaces, and varying coefficient finite mixture of regression models
  • 批准号:
    RGPIN-2015-03805
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.02万
  • 财政年份:
    2019
  • 负责人:
    Khalili, Abbas
  • 依托单位:
Statistical inference in finite mixture of regressions and mixture-of-experts models in high-dimensional spaces, and varying coefficient finite mixture of regression models
  • 批准号:
    RGPIN-2015-03805
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.02万
  • 财政年份:
    2018
  • 负责人:
    Khalili, Abbas
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
  • 批准号:
    61373035
  • 项目类别:
    面上项目
  • 资助金额:
    77.0万元
  • 批准年份:
    2013
  • 负责人:
    冯志勇
  • 依托单位: