课题基金 / 基金详情

Learning subgroups from data: selective inference and applications

Learning subgroups from data: selective inference and applications
从数据中学习子群:选择性推理和应用
批准号:
RGPIN-2021-02548
负责人:
Gao, Lucy
金额:
$1.31万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31

项目摘要

项目成果

Gao, Lucy的其他基金

相似基金

相关文献

中文摘要
翻译
像聚类和回归树这样的“子组学习”方法被用来将大型、复杂的数据集分割成尽可能同质的较小块(“子组”)。这些方法已被用于重要任务,如识别对实验药物反应不同的患者亚群,识别具有不同生物学特征的肿瘤亚群(导致有针对性的治疗策略),以及识别可能更有可能购买特定产品的客户(导致有针对性的广告策略)。在这份提案中,我概述了我的计划,即开发新的统计方法,以确定数据集中已识别的子组是否真的不同,并应用子组学习来解决称为空间填充设计的领域中的难题。一旦我们确定了数据集中的子组,自然会想知道它们是否真的不同。毕竟,如果分组学习方法确定了所有具有相同生物学特征的肿瘤亚组,或者所有具有相同购买偏好的客户亚组,那么针对这些亚组的治疗或广告策略将是巨大的时间和资源浪费。不幸的是,现有的测试两组是否不同的经典统计方法过于乐观,几乎总是声称子组不同,因为它们没有考虑到数据的双重使用:一次识别候选子组,再一次确定它们是否不同。为了解决这个问题,研究计划提出了统计方法,当测试使用聚类树和回归树识别的子组之间的方法差异时,适当地考虑到数据的双重使用。我将在一个称为空间填充设计的统计领域进一步强调子组学习方法的未开发潜力,该设计的目标是在整个空间均匀分布点。这些设计在计算机实验中得到了广泛的应用,计算机实验使用计算机模型来模拟输入变量对天气和海冰等物理系统的影响。空间填充设计也被用于环境监测网络的设计,如空气质量和水声监测网络。尽管空间在实践中往往是复杂的(例如,海岸线上的网络),但除非空间是简单的,否则很难构建空间填充设计。我提出了一种基于子组学习方法(层次聚类)的策略,可以在任意复杂的空间上构建空间填充设计。在这个研究项目中,研究生和本科生将为产生影响生物和公共卫生等领域科学进步速度的统计方法做出贡献,并成为小组学习方面的专家,这将是他们未来作为统计学家或学术界和工业界的数据科学家的职业生涯的宝贵财富。
英文摘要
"Subgroup learning" methods like clustering and regression trees are used to split large, complex data sets into smaller chunks ("subgroups") that are as homogeneous as possible. These methods have been used for important tasks like identifying subgroups of patients that respond differently to an experimental drug, identifying subgroups of tumours that have different biological profiles (leading to targeted treatment strategies), and identifying customers who might be more likely to purchase particular products (leading to targeted advertising strategies). In this proposal, I outline my plans to develop new statistical methodology to determine if identified subgroups in a data set are truly different, and to apply subgroup learning to solve a difficult problem in an area called space-filling designs.  Once we have identified subgroups in a data set, it is natural to want to know whether they are truly different. After all, if the subgroup learning method identified subgroups of tumours that all have the same biological profiles, or subgroups of customers who all have the same purchasing preferences, it would be a massive waste of time and resources to target treatment or advertising strategies to these subgroups. Unfortunately, existing classical statistical methods for testing whether two groups are different are overly optimistic and will almost always claim that the subgroups are different, because they do not account for the double use of the data: once to identify candidate subgroups, and once again to determine whether they are different. To solve this problem, the research program proposes statistical methods that properly account for the double use of the data, when testing for differences in means between subgroups identified using clustering and regression trees.  I will further highlight the untapped potential of subgroup learning methods in a statistical area called space-filling designs, which aims to evenly distribute points throughout a space. These designs are widely used in computer experiments, which use computer models to simulate the effect of input variables on physical systems like weather and sea ice. Space-filling designs are also used to design environmental monitoring networks, like air quality and underwater acoustic monitoring networks. Although spaces are often complex in practice (e.g. networks on coastlines), it is difficult to construct space-filling designs unless the space is simple. I propose a strategy based on an application of a subgroup learning method (hierarchical clustering) that can construct space-filling designs on arbitrarily complex spaces.  In this research program, graduate and undergraduate students will contribute to producing statistical methods that impact the rate of scientific advancement in areas like biology and public health, and become experts in subgroup learning, which will be a valuable asset to their future careers as statisticians or data scientists in academia and industry.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Learning subgroups from data: selective inference and applications
  • 批准号:
    RGPIN-2021-02548
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2022
  • 负责人:
    Gao, Lucy
  • 依托单位:
Learning subgroups from data: selective inference and applications
  • 批准号:
    DGECR-2021-00017
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2021
  • 负责人:
    Gao, Lucy
  • 依托单位:
Robust sparse partial least squares regression and classification
  • 批准号:
    487299-2016
  • 项目类别:
    Postgraduate Scholarships - Doctoral
  • 资助金额:
    $0.76万
  • 财政年份:
    2019
  • 负责人:
    Gao, Lucy
  • 依托单位:
Robust sparse partial least squares regression and classification
  • 批准号:
    487299-2016
  • 项目类别:
    Postgraduate Scholarships - Doctoral
  • 资助金额:
    $0.76万
  • 财政年份:
    2018
  • 负责人:
    Gao, Lucy
  • 依托单位:
海外基金