课题基金 / 基金详情

高维复杂数据分析中具有可重复性的统计学习方法研究及其应用

批准号:
72071187
项目类别:
面上项目
资助金额:
48.0 万元
负责人:
郑泽敏
依托单位:
学科分类:
管理统计理论与方法
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
郑泽敏

项目摘要

结项摘要

郑泽敏的其他基金

相似基金

相关文献

中文摘要
高维复杂数据分析是当前大数据与人工智能国家战略的重要组成部分,同时也是国际统计学及机器学习研究的热点领域,在动态定价策略、产品组合优化、平台推荐系统和社交网络推断等当代管理学热点问题中有着广泛应用。然而,由于样本中的随机噪音及高维数据的复杂结构,数据分析的结果在不同的样本下常常不具有可重复性,这给统计学习中结论的可靠性带来极大挑战。本项目拟基于在模型中引入人工控制变量构造检验统计量的思想,发展若干统计学习新方法以控制错误发现率,从而实现复杂数据分析中科学发现的可重复性。主要研究工作将从多响应回归问题中的错误发现率控制、带删失响应数据中统计学习的可重复性、多变点检测问题中具有可重复性的统计推断这三个方向展开。相关研究成果将应用于药品智慧评价与管理等实际问题。
英文摘要
High-dimensional complex data analysis is an intrinsic component of the current strategy “Big Data and Artificial Intelligence” in our country, also the forefront of statistical science and machine learning, with wide applications in hot issues of modern management such as dynamic pricing, assortment optimization, recommender systems, and social network inference. However, due to the random noises in the sample and the complex structures of high-dimensional data, the findings of data analysis may not be reproducible, which poses great challenges to the reliability of statistical learnings. In this project, we will utilize the idea of false discovery rate control through control variables to innovate several new statistical learnning methods, so that the scientific findings in complex data analysis can be reproducible. The main research contents include three aspects: false discovery rate control in multiple response regressions, reproducible statistical learning with censored responses,and reproducible change point detection in multiple change point problems. The new methods aim at real applications including intelligent evaluation and management of drugs.
高维复杂数据分析是当前大数据与人工智能国家战略的重要组成部分,同时也是国际统计学及机器学习研究的热点领域,在动态定价策略、产品组合优化、平台推荐系统和社交网络推断等当代管理学热点问题中有着广泛应用。然而,由于样本中的随机噪音及高维数据的复杂结构,数据分析的结果在不同的样本下常常不具有可重复性,这给统计学习中结论的可靠性带来极大挑战。本项目以高维统计推断方法的研究为基础,在模型中引入人工控制变量构造检验统计量的思想,发展若干统计学习新方法以控制错误发现率,解决复杂数据分析中科学发现的可重复性分析的多个关键问题,研究内容包括但不限于多响应回归问题中的错误发现率控制、带删失响应数据中统计学习的可重复性、多变点检测问题中具有可重复性的统计推断。相关研究成果发表(含在线发表)于Operations Research(1篇)、Journal of Machine Learning Research(1篇)、Journal of Business & Economic Statistic(1篇)、INFORMS Journal on Computing(3篇)、 Statistica Sinica(1篇)、Journal of Multivariate Analysis(1篇)、Computational Statistics and Data Analysis(1篇)、Communications in Mathematics and Statistics(1篇)、Statistics & Probability Letters (3篇)、Communications in Statistics - Theory and Methods (4篇)、Stat(1篇)等国际统计学、机器学习及管理优化著名期刊上,将应用于经济、金融、生物医疗以及机器学习等领域的大规模数据处理。
含有潜在因子的大规模数据统计分析及应用
  • 批准号:
    11601501
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    19.0万元
  • 批准年份:
    2016
  • 负责人:
    郑泽敏
  • 依托单位:
国内基金
海外基金