Non-uniform sampling of permutations and large scale hypothesis testing
Non-uniform sampling of permutations and large scale hypothesis testing
批准号:
1521145
负责人:
Art Owen
金额:
$39.97万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-08-01 至 2019-07-31
中文摘要
现代科学工具正在提供非常大的数据集。这在生物学中尤其正确,因为可以测量数千个基因的表达水平,甚至可以测量基因组上数百万个位置的特定DNA信息。科学家们希望将这些变量与其他测量量联系起来,特别是疾病的存在或不存在。当数以百万计的假设被调查时,其中一个可能只是偶然地与某些基因相关。人们通常坚持认为,在一次测试中观察到的相关性如此之强,以至于在2000万次测试中最多只能偶然发生一次。衡量机会相关性的常用方法是随机地对数据进行洗牌,看看强烈的影响出现的频率。如果感兴趣的事件是2千万分之一的结果,我们通常需要大约10倍的随机洗牌来确定。这个提议是关于寻找更有效的随机洗牌策略,以更少的洗牌次数获得期望的答案。目标是以更少的计算和更高的可靠性找到重要的生物变量。找到重要的基因是后续工作的第一步,后续工作包括挖掘文献和进行实验,以了解这些基因的作用,并确定它们之间的关系是否有用。部分工作还将涉及调整其他测量因素或其他可能使观察到的相关性具有误导性的因素。用于发现和测量罕见和不寻常结果的新数学方法也可用于工业问题,其中罕见现象是通过计算机模拟测量的异常有效的产品设计。测试一个基因或一组基因是否与表型(疾病、身高等)或治疗(饮食、药物等)相关的通常方法是进行排列测试。从n个数据点,有多达n个!排列运行。通常这个数量的排列超出了我们的预算,我们也从这些排列中抽样。如果我们计算检验统计量M次,对原始数据计算一次,对M-1个排列各计算一次,那么我们可能得到的最小p值是1/M。也就是说,为了获得目标p值,我们必须至少计算1/p次统计数据。全基因组关联研究的标准门槛转化为至少2000万次计算。要在排列测试中有足够的功率,需要更多的10/p计算。当表型/处理是二元的,排列试验减少到抽样与替换。这个项目使用排列或组合的非均匀抽样。主要方法是利用混合分量概率作为控制变量,从混合方案中进行重要抽样。马尔可夫链蒙特卡罗方法将进行研究。
英文摘要
Modern scientific tools are delivering very large data sets. This is especially true in biology where expression levels for thousands of genes or even the specific DNA information at millions of locations on the genome can be measured. Scientists would like to correlate these variables with other measured quantities, especially the presence or absence of a disease. When millions of hypotheses are investigated, it is possible that one of them will correlate with some genes just by chance. It is common to insist that the observed correlation for one test be so strong that it would happen by chance at most once in 20 million tries. The usual way to measure chance correlations is to shuffle the data at random and see how often a strong effect appears. If the event of interest is a one in 20 million outcome we usually need about ten times that many random shuffles to be sure. This proposal is about finding more efficient random shuffling strategies to get a desired answer with fewer shuffles. The goal is to find important biological variables with much less computation and greater reliability. Finding the important genes is a first step for followup work that includes mining the literature and running experiments to understand the role of those genes and determine whether their relationship is useful or not. Part of the work will also involve adjusting for other factors measured or otherwise that could make the observed correlations misleading. New mathematical methods for finding and measuring rare and unusual outcomes can also be used in industrial problems where the rare phenomenon is an unusually effective product design as measured by computer simulations.The usual way to test whether a gene or a gene set is associated with a phenotype (disease, height, etc.) or a treatment (diet, medicines, etc.) is to run a permutation test. From n data points, there are as many as n! permutations to run. Usually this amount of permutations is beyond our budget and we sample from the permutations as well. If we compute our test statistic M times, once on the original data and once for each of M-1 permutations, then the smallest p value we can possibly get is 1/M. That is, to attain a target p value we have to compute our statistic at least 1/p times. The standard threshold for genome wide association studies translates into a bare minimum of 20,000,000 computations. To have adequate power in a permutation test requires more like 10/p computations. When the phenotype/treatment is binary, the permutation test reduces to sampling with replacement. This project uses non-uniform sampling of permutations or combinations. The main method is importance sampling from mixtures of proposals using the mixture component probabilities as control variates. Markov chain Monte Carlo methods will be investigated.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Randomized quasi-Monte Carlo sampling for scientific computing
-
批准号:2152780
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2022
-
负责人:Art Owen
-
依托单位:
BIGDATA: F: Computationally Efficient Algorithms for Large-Scale Crossed Random Effects Models
-
批准号:1837931
-
项目类别:Standard Grant
-
资助金额:$80.0万
-
财政年份:2018
-
负责人:Art Owen
-
依托单位:
Monte Carlo and Quasi-Monte Carlo Methods for Statistics
-
批准号:1407397
-
项目类别:Continuing Grant
-
资助金额:$22.5万
-
财政年份:2014
-
负责人:Art Owen
-
依托单位:
MCQMC 2014 Travel Support
-
批准号:1357690
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2014
-
负责人:Art Owen
-
依托单位:
MCQMC 2012
-
批准号:1135257
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:2011
-
负责人:Art Owen
-
依托单位:
Monte Carlo and Quasi-Monte Carlo Methods for Statistics
-
批准号:0906056
-
项目类别:Continuing Grant
-
资助金额:$65.98万
-
财政年份:2009
-
负责人:Art Owen
-
依托单位:
Travel support for MCQMC July 2008, Montreal, Canada
-
批准号:0805890
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2008
-
负责人:Art Owen
-
依托单位:
Monte Carlo and Quasi-Monte Carlo Methods for Statistics
-
批准号:0604939
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2006
-
负责人:Art Owen
-
依托单位:
Statistical Integration and Approximation
-
批准号:0306612
-
项目类别:Continuing Grant
-
资助金额:$44.33万
-
财政年份:2003
-
负责人:Art Owen
-
依托单位:
Statistical Numerics
-
批准号:0072445
-
项目类别:Continuing Grant
-
资助金额:$21.0万
-
财政年份:2000
-
负责人:Art Owen
-
依托单位:
Statistical Integration and Approximation
-
批准号:9704495
-
项目类别:Continuing Grant
-
资助金额:$18.0万
-
财政年份:1997
-
负责人:Art Owen
-
依托单位:
Mathematical Sciences Computing Research Environments
-
批准号:9508275
-
项目类别:Standard Grant
-
资助金额:$4.95万
-
财政年份:1995
-
负责人:Art Owen
-
依托单位:
Mathematical Sciences: Investigations Into Computationally Intensive Statistical Methods
-
批准号:9404594
-
项目类别:Standard Grant
-
资助金额:$6.0万
-
财政年份:1994
-
负责人:Art Owen
-
依托单位:
U.S. - Australian Cooperative Research: Computer Intensive Statistics
-
批准号:8913333
-
项目类别:Standard Grant
-
资助金额:$1.58万
-
财政年份:1991
-
负责人:Art Owen
-
依托单位:
Mathematical Sciences: Sampling Theory for Computer Experiments
-
批准号:9011074
-
项目类别:Continuing Grant
-
资助金额:$8.93万
-
财政年份:1990
-
负责人:Art Owen
-
依托单位:
国内基金
海外基金
基于Riemann-Hilbert方法的相关问题研究
-
批准号:11026205
-
项目类别:数学天元基金项目
-
资助金额:3.0万元
-
批准年份:2010
-
负责人:周建荣
-
依托单位:
微分遍历理论和廖山涛的一些方法的应用
-
批准号:10671006
-
项目类别:面上项目
-
资助金额:21.0万元
-
批准年份:2006
-
负责人:孙文祥
-
依托单位: