课题基金 / 基金详情

Statistical learning algorithms for high-dimensional non-normally distributed data

Statistical learning algorithms for high-dimensional non-normally distributed data
高维非正态分布数据的统计学习算法
批准号:
RGPIN-2018-06787
负责人:
Shaikh, Mateen
金额:
$1.17万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Shaikh, Mateen的其他基金

相似基金

相关文献

中文摘要
翻译
反映各种数据的计算方法正在不断改进。该提案提出了用统计学习技术解决建模和发现数据模式问题的三条主线。*研究的第一条主线是二进制数据库。从跟踪客户购买内容的收据记录到表明事故涉及的因素的观察,二进制数据库都很常见,而且可能非常大。出于各种原因,总结这些数据库中的关联是一项有用的任务。存在许多可能的关联,将它们相互比较是重要的。对这些关联进行数值比较特别有价值,因为这可以由计算机在大范围内自动进行。然而,数字摘要的选择很重要。这项建议提出了改进二进制数据中关联的数字摘要以阐明模式的方法。其中一种方法是在一些二元变量实际上是分类变量的元素时如何汇总数据,另一种方法是根据数据的分布来考虑这些值有多值得注意。*提案的第二条线索涉及模型的复杂性。尽管非常复杂的模型可以准确地对某些数据进行建模,但这是不可取的,原因有很多,包括可解释性、健壮性和计算挑战。通过考虑某些参数(定义模型的量)被约束为与模型的其他参数相同时,可以简化一些复杂的模型。这样无论参数代表什么,都可以减少计算机所需的估计次数,并使模型更易于解释。这项建议建议探索一种最近提出的为各种统计模型发现这些约束的方法。*这项建议的最后一条主线解决了在对通常被认为是“连续”的变量建模时所作假设的现实问题。这些数据通常被建模为真正连续的,遵循特定的分布(正态分布),或者两者兼而有之。在这个思路中,考虑了更灵活的假设,并适应了表示连续变量的数据实际上只知道有限精度的情况,这可能会影响结果。探索将确定在哪些情况下这种有限的精度很重要,以及当考虑到有限的精度和不那么严格的假设时,答案的准确性有多高。*所有这些问题都将随着高素质人员开发和应用新技能而得到解决,这些技能经过分析现实数据(有时是不方便的)和大数据的培训。这是一种技能,已被确定为加拿大国内的“人才缺口”,并将在本提案中得到解决。
英文摘要
Computational methods to reflect a variety of data are continuing to improve. This proposal suggests three main threads of addressing issues with modelling and discovering patterns in data with statistical learning techniques.******The first thread of research considers binary data bases. From records of receipts keeping track of what customers purchase to observations indicating the factors involved in an accident, binary data bases are both common, and can be very large. Summarizing associations in these data bases is a useful task for a variety of reasons. Many possible associations exist and comparing them to each other is important. Numerically comparing these associations is particularly valuable as this can be automated by the computer in large scales. However, the choice of numerical summary is important. This proposal suggests methods of improving numerical summaries of associations in binary data to elucidate patterns. One of these methods is on how to summarize data when some of the binary variables are actually elements of a categorical variable, and another is to consider how noteworthy these values are in light of the distribution of the data. ******A second thread of the proposal addresses the complexity of models. Although very complex models can accurately model some data, this is undesirable for a variety of reasons including interpretability, robustness, and computational challenges. Some complex models can be simplified by considering when certain parameters, quantities which define a model, are constrained to be the same as other parameters of the model. This relates whatever the parameters represent, reduces the number of estimates the computer requires, and make the model easier to interpret. This proposal suggests explores a recently proposed method of discovering these constraints for a variety of statistical models.****** The final thread of this proposal addresses the realistic issue of the assumptions made when modelling what are often considered to be "continuous" variables. These data are often modeled as truly continuous, following a particular distribution (the normal distribution), or both. In this thread, more flexible assumptions are considered and accommodates the situation that data representing continuous variables are actually only known up to a limited precision which can influence results. The exploration will determine in which scenarios this limited precision matters and how accurate answers are when accounting for the limited precision and less stringent assumptions. ******All of these issues will be addressed as highly qualified personnel develop and apply new skills, trained in the analysis of realistic, sometimes inconvenient, and big data. This is a skillset that has been identified as a "talent gap" within Canada and will be addressed with this proposal.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical learning algorithms for high-dimensional non-normally distributed data
  • 批准号:
    RGPIN-2018-06787
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.17万
  • 财政年份:
    2022
  • 负责人:
    Shaikh, Mateen
  • 依托单位:
Statistical learning algorithms for high-dimensional non-normally distributed data
  • 批准号:
    RGPIN-2018-06787
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.17万
  • 财政年份:
    2021
  • 负责人:
    Shaikh, Mateen
  • 依托单位:
Statistical learning algorithms for high-dimensional non-normally distributed data
  • 批准号:
    RGPIN-2018-06787
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.17万
  • 财政年份:
    2020
  • 负责人:
    Shaikh, Mateen
  • 依托单位:
Statistical learning algorithms for high-dimensional non-normally distributed data
  • 批准号:
    RGPIN-2018-06787
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.17万
  • 财政年份:
    2019
  • 负责人:
    Shaikh, Mateen
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: