课题基金 / 基金详情

Sparsity, thresholding and regularization in data science

Sparsity, thresholding and regularization in data science
数据科学中的稀疏性、阈值化和正则化
批准号:
RGPIN-2022-04531
负责人:
DiazRodriguez, Jairo
金额:
$1.38万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

DiazRodriguez, Jairo的其他基金

相似基金

相关文献

中文摘要
翻译
数据科学为现实生活中的决策方式带来了突破,如欺诈检测、医疗保健、定向广告、网站推荐、语音识别等。经典的和新的统计学习技术的实际应用是这些进步的共同点。然而,大多数这样的应用程序仍然被其创建者误解,并且解决方案主要是在试验和错误中实现的。机器学习中最有用的技术之一是正则化。它有助于处理过拟合问题,但也会在优化算法的求解中施加结构。阈值估计器是一组特殊的正则化估计器,它在解中施加稀疏结构。稀疏性假设只有少数协变量组成模型来解释给定的响应。例如,只有少数基因与解释特定疾病相关。此外,稀疏性可以赋予结果可解释性或物理意义。本提案的目标是通过使用稀疏正则化和阈值估计器来发展解决和理解机器和统计学习模型的理论和创新方法。该提案包括以下三个研究方向。首先,高维数据经常出现在计量经济学、机器学习、神经科学和社会科学中。我将把我以前在阈值估计方面的工作扩展到高维数据的其他方法。我将对涉及社会科学分类数据的应用感兴趣。第二,我对间接传播疾病的风险预测很感兴趣。这可以看作是一个高维层析反问题。目标是通过施加稀疏的总变异空间结构,通过使用GPS动物运动等非侵入性测量方法重建高和低疾病风险区域,对一个地区进行流行病学断层扫描。由此产生的方法将在完整的数据科学框架中实施。最后,我建议使用阈值估计器将稀疏结构强加到深度学习方法中。目前的深度自动编码器倾向于强迫神经网络的架构。我建议使用阈值正则化器来共同估计网络结构。我还提出了一个新的基于L1正则化的dropout框架。我建议对正则化参数进行随机选择,而不是随机丢弃单元。这些涉及稀疏性的方法学可能导致对方法产生更好的解释,并可能促进数学性质的推导。本研究项目的成功将对高维数据、机器学习和大数据的理解做出重大贡献,并将促进可解释稀疏正则化在自然科学、社会科学和工程等诸多领域的应用。
英文摘要
Data science has brought a breakthrough in the way decisions are made in real life problems such as fraud detection, healthcare, targeted advertising, website recommendations, speech recognition, among others. Practical implementations of classic and new statistical earning techniques are the common denominator of such advances. However, most of such applications are still misunderstood by its creators, and solutions are mainly implemented out of trial and error. One of the most useful techniques in machine learning is regularization. It helps to cope with overfitting problems but also impose structures in the solution of the optimization algorithms. Thresholding estimators are a particular set of regularization estimators that impose sparse structure in the solution. Sparsity assumes that only a few covariates compose the model to explain a given response. For instance, just a few genes are relevant to explain a given disease. Moreover, sparsity can give interpretability or physical meaning to the result. The objective of this proposal is to develop theory and innovative methodologies for solving and understanding machine and statistical learning models by using sparse regularization and thresholding estimators. The proposal consists of the following three lines of research. First, high dimensional data routinely arise in econometrics, machine learning, neuroscience, and social science. I will extend my previous work in thresholding estimators to other methodologies for high dimensional data. I will be interested in applications involving categorical data for social sciences. Second, I am interested in the prediction of the risk of indirectly transmitted diseases. This can be seen as a high dimensional tomographic inverse problem. The objective is to perform an epidemiologic tomography of a region by reconstructing the areas of high and low disease risk using non-invasive measurements such as GPS animal movements, by imposing a sparse total variation spatial structure. The resulting methodology will be implemented in a full data science framework. Finally, I propose to use thresholding estimators to impose sparse structures into Deep Learning methodologies. Current deep autoencoders tend to force the architecture of a neural network. I propose to impose thresholding regularizers to jointly estimate the network architecture. I also propose a new dropout framework based on L1 regularization. Instead of randomly dropping units, I propose to perform a random selection on the regularization parameter. These methodologies involving sparsity might lead to produce better interpretation of the methods and might facilitate the derivation of mathematical properties. The success of this research program will have great contribution to the understanding of high dimensional data, machine learning, and big data, and will prompt the applications of interpretable sparse regularization in many fields in the natural sciences, social sciences, and engineering.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Sparsity, thresholding and regularization in data science
  • 批准号:
    DGECR-2022-00453
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2022
  • 负责人:
    DiazRodriguez, Jairo
  • 依托单位:
海外基金