课题基金 / 基金详情

Novel p-Value Based Multiple Testing Methods for Variable Selection with False Discovery Rate Control

Novel p-Value Based Multiple Testing Methods for Variable Selection with False Discovery Rate Control
基于 p 值的新颖变量选择多重测试方法以及错误发现率控制
批准号:
2210687
负责人:
Sanat Sarkar
金额:
$26.97万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2025-06-30

项目摘要

项目成果

Sanat Sarkar的其他基金

相似基金

相关文献

中文摘要
翻译
多重检验是现代科学调查中遇到的最常见的统计挑战之一。该项目旨在解决长期存在的多种测试方法的应用问题。其中一个问题是在一大堆变量中发现那些对感兴趣的结果有重要影响的变量。由于变量之间未知的相互依赖关系,标准的多种测试方法不适用就是这样一个问题。正在开发的方法旨在提供新的方法来发现重要的变量,无论这些变量如何相互依赖,并保证平均而言,只有一小部分受控的不重要变量最终成为错误发现。一个例子是在识别基因变异方面,在成千上万的遗传变异中,这些变异可以影响某种疾病。新方法可以帮助识别与治疗干预相对更相关的基因。这些方法开发背后的基本理论和方法论思想将扩展到解决在其他实验环境中使用多种测试方法的类似问题。该项目将开展的研究将纳入课程,有利于本科生和研究生的培养。本研究项目侧重于解决与多次测试相关的重要理论和方法问题。例如,在具有高斯噪声的多元线性回归背景下的特征/变量选择在数据科学中占有重要地位,在科学研究中是一个普遍存在的统计框架,它往往被框架化为多重测试问题。基于p值的多重检验方法将是最理想的,无论考虑什么错误率来控制错误发现的重要解释变量,完全捕获解释变量的相关矩阵而不失去对错误率的控制。不幸的是,这种方法还没有在非渐近的环境中得到发展。类似地,对于具有非对角相关矩阵的多元高斯均值的同时检验的相关问题,在错误率的控制下,基于p值的多重检验方法在不失去对错误率的控制的情况下完全捕获相关信息的方法在文献中基本上是空白的。这些挑战将通过将两个关于多重推理的开创性想法交叉授粉的研究来应对:1)使用基于p值的多重测试方法来控制错误发现;以及2)在线性回归设置中使用设计矩阵的仿冒来选择变量。具体地说,该项目旨在开发新的基于p值的错误发现率和其他控制多种测试方法的强大错误率,用于1)在低维和高维环境中带有高斯噪声的多元线性回归中的变量选择;以及2)同时测试具有一般非对角线协方差矩阵的多变量高斯均值。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Multiple testing is one of the most common statistical challenges encountered in modern scientific investigations. This project aims at resolving some longstanding issues with application of multiple testing methods. One of these issues arises in the context of discovering, among a large collection of variables, those that are important influences on an outcome of interest. Inapplicability of standard multiple testing methods due to the unknown interdependency of the variables is such an issue. The methods under development aim to provide new approaches to discovering important variables no matter how the variables depend on each other, with the guarantee that, on average, only a small, controlled fraction of unimportant variables end up as false discoveries. An example application is in the identification of genetic variants which, among many thousands of them, can influence a certain disease. The new methods can aid in identifying genes as being relatively more relevant for therapeutic intervention. The fundamental theoretical and methodological ideas behind the development of these methods will be extended towards resolving similar issues with multiple testing methods in other experimental settings as well. The research to be carried out in the project will be incorporated into courses, benefiting the training of undergraduates and graduate students.This research project is focused on addressing important theoretical and methodological issues related to multiple testing. For instance, feature/variable selection under the setting of multiple linear regression with Gaussian noise, which plays an important role in data science and is a ubiquitous statistical framework in scientific investigations, is often framed as a multiple testing problem. A p-value based multiple testing method, irrespective of what error rate is being considered to control the falsely discovered important explanatory variables, capturing the correlation matrix of the explanatory variables in full without losing control over the error rate, would be most ideal. Unfortunately, such methods are yet to be developed in a non-asymptotic setting. Similarly, for the related problem of simultaneous testing of multivariate Gaussian means with non-diagonal correlation matrix, subject to a control of an error rate, a p-value based multiple testing method fully capturing the correlation information without losing control over that rate is largely absent from the literature. The challenges will be met by research that cross-fertilizes two seminal ideas on multiple inference: 1) the use of p-value based multiple testing methods to control false discoveries; and 2) the use of the knockoff of the design matrix for variable selection in linear regression settings. Concretely, the project aims at developing novel p-value based false discovery rate and other powerful error rates controlling multiple testing methods for 1) variable selection in multiple linear regression with Gaussian noise, both in low- and high-dimensional settings; and 2) simultaneous testing of multivariate Gaussian means with a general non-diagonal covariance matrix.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: New Directions for Research on Some Large-Scale Multiple Testing Problems
  • 批准号:
    1309273
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $12.66万
  • 财政年份:
    2013
  • 负责人:
    Sanat Sarkar
  • 依托单位:
Collaborative Research: Constructing New Multiple Testing Methods
  • 批准号:
    1006344
  • 项目类别:
    Standard Grant
  • 资助金额:
    $16.77万
  • 财政年份:
    2010
  • 负责人:
    Sanat Sarkar
  • 依托单位:
Multiple Testing: Further Development Of Theory And Methodology
  • 批准号:
    0603868
  • 项目类别:
    Standard Grant
  • 资助金额:
    $16.99万
  • 财政年份:
    2006
  • 负责人:
    Sanat Sarkar
  • 依托单位:
New Problems in Multiple Hypotheses Testing
  • 批准号:
    0306366
  • 项目类别:
    Standard Grant
  • 资助金额:
    $23.4万
  • 财政年份:
    2003
  • 负责人:
    Sanat Sarkar
  • 依托单位:
国内基金
海外基金
基于时间序列间分位相依性(quantile dependence)的风险值(Value-at-Risk)预测模型研究
  • 批准号:
    71903144
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    17.0万元
  • 批准年份:
    2019
  • 负责人:
    张申
  • 依托单位: