课题基金 / 基金详情

Novel p-Value Based Multiple Testing Methods for Variable Selection with False Discovery Rate Control

Novel p-Value Based Multiple Testing Methods for Variable Selection with False Discovery Rate Control
基于 p 值的新颖变量选择多重测试方法以及错误发现率控制
批准号:
2210687
负责人:
Sanat Sarkar
金额:
$26.97万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2025-06-30

项目摘要

项目成果

Sanat Sarkar的其他基金

相似基金

相关文献

中文摘要
翻译
多重检验是现代科学调查中最常见的统计挑战之一。本项目旨在通过应用多种测试方法解决一些长期存在的问题。其中一个问题出现在在大量变量中发现那些对感兴趣的结果有重要影响的变量的背景下。由于变量的未知相互依赖关系,标准的多重测试方法不适用就是这样一个问题。正在开发的方法旨在提供发现重要变量的新方法,而不管这些变量是如何相互依赖的,并保证平均而言,只有一小部分可控的不重要变量最终成为错误的发现。一个例子应用是在成千上万的基因变异中识别能够影响某种疾病的基因变异。新方法可以帮助识别与治疗干预相对更相关的基因。这些方法发展背后的基本理论和方法思想将扩展到解决其他实验环境中多种测试方法的类似问题。本项目所进行的研究将纳入课程,有利于本科生和研究生的培养。本研究项目的重点是解决与多重测试相关的重要理论和方法问题。例如,带高斯噪声的多元线性回归背景下的特征/变量选择在数据科学中起着重要作用,是科学研究中普遍存在的统计框架,但往往被定义为一个多重检验问题。基于p值的多重检验方法,无论考虑何种错误率来控制错误发现的重要解释变量,在不失去对错误率控制的情况下,充分捕获解释变量的相关矩阵,将是最理想的。不幸的是,这种方法尚未在非渐近环境中得到发展。同样,对于非对角相关矩阵的多元高斯均值同时检验的相关问题,在错误率控制的情况下,基于p值的多重检验方法在不失去对错误率控制的情况下充分捕获相关信息,在文献中基本没有。这些挑战将通过对多重推理的两个开创性思想进行交叉施肥的研究来解决:1)使用基于p值的多重测试方法来控制错误发现;2)在线性回归设置中使用设计矩阵的仿制品进行变量选择。具体而言,该项目旨在开发新的基于p值的错误发现率和其他强大的错误率控制多种测试方法,用于1)低维和高维设置下高斯噪声多元线性回归中的变量选择;2)用一般非对角协方差矩阵同时检验多元高斯均值。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Multiple testing is one of the most common statistical challenges encountered in modern scientific investigations. This project aims at resolving some longstanding issues with application of multiple testing methods. One of these issues arises in the context of discovering, among a large collection of variables, those that are important influences on an outcome of interest. Inapplicability of standard multiple testing methods due to the unknown interdependency of the variables is such an issue. The methods under development aim to provide new approaches to discovering important variables no matter how the variables depend on each other, with the guarantee that, on average, only a small, controlled fraction of unimportant variables end up as false discoveries. An example application is in the identification of genetic variants which, among many thousands of them, can influence a certain disease. The new methods can aid in identifying genes as being relatively more relevant for therapeutic intervention. The fundamental theoretical and methodological ideas behind the development of these methods will be extended towards resolving similar issues with multiple testing methods in other experimental settings as well. The research to be carried out in the project will be incorporated into courses, benefiting the training of undergraduates and graduate students.This research project is focused on addressing important theoretical and methodological issues related to multiple testing. For instance, feature/variable selection under the setting of multiple linear regression with Gaussian noise, which plays an important role in data science and is a ubiquitous statistical framework in scientific investigations, is often framed as a multiple testing problem. A p-value based multiple testing method, irrespective of what error rate is being considered to control the falsely discovered important explanatory variables, capturing the correlation matrix of the explanatory variables in full without losing control over the error rate, would be most ideal. Unfortunately, such methods are yet to be developed in a non-asymptotic setting. Similarly, for the related problem of simultaneous testing of multivariate Gaussian means with non-diagonal correlation matrix, subject to a control of an error rate, a p-value based multiple testing method fully capturing the correlation information without losing control over that rate is largely absent from the literature. The challenges will be met by research that cross-fertilizes two seminal ideas on multiple inference: 1) the use of p-value based multiple testing methods to control false discoveries; and 2) the use of the knockoff of the design matrix for variable selection in linear regression settings. Concretely, the project aims at developing novel p-value based false discovery rate and other powerful error rates controlling multiple testing methods for 1) variable selection in multiple linear regression with Gaussian noise, both in low- and high-dimensional settings; and 2) simultaneous testing of multivariate Gaussian means with a general non-diagonal covariance matrix.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: New Directions for Research on Some Large-Scale Multiple Testing Problems
  • 批准号:
    1309273
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $12.66万
  • 财政年份:
    2013
  • 负责人:
    Sanat Sarkar
  • 依托单位:
Collaborative Research: Constructing New Multiple Testing Methods
  • 批准号:
    1006344
  • 项目类别:
    Standard Grant
  • 资助金额:
    $16.77万
  • 财政年份:
    2010
  • 负责人:
    Sanat Sarkar
  • 依托单位:
Multiple Testing: Further Development Of Theory And Methodology
  • 批准号:
    0603868
  • 项目类别:
    Standard Grant
  • 资助金额:
    $16.99万
  • 财政年份:
    2006
  • 负责人:
    Sanat Sarkar
  • 依托单位:
New Problems in Multiple Hypotheses Testing
  • 批准号:
    0306366
  • 项目类别:
    Standard Grant
  • 资助金额:
    $23.4万
  • 财政年份:
    2003
  • 负责人:
    Sanat Sarkar
  • 依托单位:
国内基金
海外基金
基于时间序列间分位相依性(quantile dependence)的风险值(Value-at-Risk)预测模型研究
  • 批准号:
    71903144
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    17.0万元
  • 批准年份:
    2019
  • 负责人:
    张申
  • 依托单位: