课题基金 / 基金详情

Non-uniform sampling of permutations and large scale hypothesis testing

Non-uniform sampling of permutations and large scale hypothesis testing
排列的非均匀采样和大规模假设检验
批准号:
1521145
负责人:
Art Owen
金额:
$39.97万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-08-01 至 2019-07-31

项目摘要

项目成果

Art Owen的其他基金

相似基金

相关文献

中文摘要
翻译
现代科学工具正在提供非常大的数据集。这在生物学中尤其如此,在生物学中,可以测量数千个基因的表达水平,甚至可以测量基因组上数百万个位置的特定DNA信息。科学家们希望将这些变量与其他测量的数量相关联,特别是疾病的存在或不存在。当数以百万计的假说被研究时,其中一个可能只是偶然地与某些基因相关。人们通常坚持认为,一次测试的观察到的相关性如此之强,以至于它最多在2000万次尝试中发生一次。衡量机率相关性的通常方法是随机调整数据,看看强效应出现的频率有多高。如果感兴趣的事件是2000万个结果中的一个,我们通常需要大约十倍的随机洗牌才能确定。这个建议是关于寻找更有效的随机洗牌策略,以更少的洗牌获得所需的答案。目标是用更少的计算和更高的可靠性找到重要的生物变量。找到重要的基因是后续工作的第一步,包括挖掘文献和进行实验,以了解这些基因的作用,并确定它们之间的关系是否有用。这项工作的一部分还将涉及对其他测量或其他因素进行调整,这些因素可能会使观察到的相关性具有误导性。发现和测量罕见和不寻常结果的新的数学方法也可以用于工业问题,在工业问题中,罕见现象是通过计算机模拟测量的异常有效的产品设计。测试基因或基因集是否与表型(疾病、身高等)有关的通常方法。或治疗(饮食、药物等)就是进行一项置换测试。从n个数据点中,有多达n!要运行的排列。通常,这样的排列数量超出了我们的预算,我们也会从排列中进行采样。如果我们计算测试统计量M次,一次是在原始数据上,一次是针对M-1个排列,那么我们可能得到的最小p值是1/M。也就是说,要获得目标p值,我们必须至少计算我们的统计量1/p次。全基因组关联研究的标准门槛转化为最少2000万次计算。为了在置换测试中具有足够的能力,需要进行更多类似于10/p的计算。当表型/处理为二元时,排列测试简化为采样和替换。该项目使用排列或组合的非均匀抽样。主要方法是使用混合成分概率作为控制变量,从混合方案中进行重要抽样。将研究马尔可夫链蒙特卡罗方法。
英文摘要
Modern scientific tools are delivering very large data sets. This is especially true in biology where expression levels for thousands of genes or even the specific DNA information at millions of locations on the genome can be measured. Scientists would like to correlate these variables with other measured quantities, especially the presence or absence of a disease. When millions of hypotheses are investigated, it is possible that one of them will correlate with some genes just by chance. It is common to insist that the observed correlation for one test be so strong that it would happen by chance at most once in 20 million tries. The usual way to measure chance correlations is to shuffle the data at random and see how often a strong effect appears. If the event of interest is a one in 20 million outcome we usually need about ten times that many random shuffles to be sure. This proposal is about finding more efficient random shuffling strategies to get a desired answer with fewer shuffles. The goal is to find important biological variables with much less computation and greater reliability. Finding the important genes is a first step for followup work that includes mining the literature and running experiments to understand the role of those genes and determine whether their relationship is useful or not. Part of the work will also involve adjusting for other factors measured or otherwise that could make the observed correlations misleading. New mathematical methods for finding and measuring rare and unusual outcomes can also be used in industrial problems where the rare phenomenon is an unusually effective product design as measured by computer simulations.The usual way to test whether a gene or a gene set is associated with a phenotype (disease, height, etc.) or a treatment (diet, medicines, etc.) is to run a permutation test. From n data points, there are as many as n! permutations to run. Usually this amount of permutations is beyond our budget and we sample from the permutations as well. If we compute our test statistic M times, once on the original data and once for each of M-1 permutations, then the smallest p value we can possibly get is 1/M. That is, to attain a target p value we have to compute our statistic at least 1/p times. The standard threshold for genome wide association studies translates into a bare minimum of 20,000,000 computations. To have adequate power in a permutation test requires more like 10/p computations. When the phenotype/treatment is binary, the permutation test reduces to sampling with replacement. This project uses non-uniform sampling of permutations or combinations. The main method is importance sampling from mixtures of proposals using the mixture component probabilities as control variates. Markov chain Monte Carlo methods will be investigated.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Randomized quasi-Monte Carlo sampling for scientific computing
  • 批准号:
    2152780
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2022
  • 负责人:
    Art Owen
  • 依托单位:
BIGDATA: F: Computationally Efficient Algorithms for Large-Scale Crossed Random Effects Models
  • 批准号:
    1837931
  • 项目类别:
    Standard Grant
  • 资助金额:
    $80.0万
  • 财政年份:
    2018
  • 负责人:
    Art Owen
  • 依托单位:
Monte Carlo and Quasi-Monte Carlo Methods for Statistics
  • 批准号:
    1407397
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $22.5万
  • 财政年份:
    2014
  • 负责人:
    Art Owen
  • 依托单位:
MCQMC 2014 Travel Support
  • 批准号:
    1357690
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2014
  • 负责人:
    Art Owen
  • 依托单位:
国内基金
海外基金
基于Riemann-Hilbert方法的相关问题研究
  • 批准号:
    11026205
  • 项目类别:
    数学天元基金项目
  • 资助金额:
    3.0万元
  • 批准年份:
    2010
  • 负责人:
    周建荣
  • 依托单位:
微分遍历理论和廖山涛的一些方法的应用
  • 批准号:
    10671006
  • 项目类别:
    面上项目
  • 资助金额:
    21.0万元
  • 批准年份:
    2006
  • 负责人:
    孙文祥
  • 依托单位: