Statistical Methods for High Dimensional Discrete Data
Statistical Methods for High Dimensional Discrete Data
批准号:
1007801
负责人:
Naomi Altman
金额:
$20.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-06-01 至 2015-05-31
中文摘要
高维计数和二进制数据现在在许多领域都很常见,包括机器学习、成像和市场营销。在高通量生物学中,产生计数和分类数据的超高通量测序技术正在取代微阵列和其他“组学”测量设备。这些测量设备的输出是针对每个样本数万个响应的每个基因或其他生物亚单位的计数,或者针对每个样本可能数百万个响应的单核苷酸多态(SNPs)等特征的存在/不存在。卫星图像、医学扫描、监测设备和其他非常高维的测量设备上的特征也可以获得类似的数据。研究人员将把为连续(主要是正态分布)数据开发的高度多变量和多重测试方法扩展到离散数据。将在四个领域开发新的方法:a)分析离散数据的分布差异,以适应使用具有超离散度和贝叶斯或经验贝叶斯收缩的广义线性混合模型的复杂实验设计。B)考虑到离散预报器的误差结构,对离散数据环境中的样本和变量进行监督聚类的方法。C)经典的、充分的降维方法,如离散数据的典型相关和分段逆回归。D)在多个测试中的概念和方法的扩展,例如错误发现率估计到离散设置,其中使用条件混合建模来自独立或弱依赖测试的p值可能具有不同的零分布。这些方法将在基因组学和成像数据上进行测试。非常多元的数据现在已经成为细胞生物学、营销、医学和卫星成像、气象学、流行病学、欺诈检测和癌症研究等领域的标准数据。这些数据可能包括对样本中每一项的数千或数百万次测量。例如,基因分型服务为个人提供关于他们细胞中数十万个基因变异的信息,零售商数据库可能有关于连锁店中每一家商店数万件商品的销售信息。其中许多数据是以计数的形式(如库存中每种类型的物品数量、编码特定蛋白质的mRNA分子的数量)或以类别的形式(如开/关、存在/不存在或AA、AA或AA)的形式。血压和体温等高度多变量连续测量的方法已经很成熟,但不能直接应用于计数和分类数据。调查员将开发统计方法和软件,以改进对计数和分类数据的分析和汇总。提出了四个主要的研究领域:A)确定变量是否与组之间的差异有关的统计模型和检验;B)预测或分类组成员的统计方法;C)用更小的派生变量集汇总数据的方法,以保持全部数据的预测能力;D)估计错误率的多重比较方法。例如,在一项与转移性癌症和非转移性癌症相关的基因研究中,这些方法可以用来确定哪些基因在发生或没有进展到转移的肿瘤中表达不同,选择一组较小的基因作为诊断工具,然后提供便于临床医生解释的摘要。在对机械零件的应力研究中,可以使用施加应力之前和施加应力期间零件扫描的像素来确定零件可能失效的精确位置以及在低应力和高应力下失效的零件之间的扫描特征之间的差异。在安装大量模型或进行测试的研究中,有必要容忍一小部分误差。为连续数据开发的多重测试的概念和方法将被扩展,以帮助估计和控制使用计数和分类数据的错误结论的数量。
英文摘要
Very high dimensional count and binary data are now common in many fields including machine learning, imaging and marketing. In high-throughput biology, ultra-high thoughput sequencing technologies which produce count and categorical data are displacing microarrays and other "omics" measurement devices. The output of these measurement devices are counts per gene or other biological subunit for tens of thousands of responses per sample, or presence/absence for features such as single nucleotide polymorphisms (SNPs), for possibly millions of responses per sample. Similar data can be derived on for features on satellite images, medical scans, monitoring devices and other very high dimensional measurement devices. The investigator will extend highly multivariate and multiple testing methods developed for continuous (primarily normally distributed) data to discrete data. New methods will be developed in four areas: A) analyses for differences in distribution for discrete data that can accommodate complex experimental designs using generalized linear mixed models with overdispersion and Bayesian or empirical Bayes shrinkage. B) methods for supervised clustering of samples and variables in the discrete data setting taking into account the error structure of the discrete predictors. C) classical and sufficient dimensions reduction methods such as canonical correlation and sliced inverse regression for discrete data. D) extension of concepts and methods in multiple testing, such as false discovery rate estimation to the discrete setting in which the p-values from independent or weakly dependent tests may have different null distributions using conditional mixture modeling. The methods will be tested on genomics and imaging data.Very highly multivariate data are now the norm in fields as diverse as cell biology, marketing, medical and satellite imaging, meteorology, epidemiology, fraud detection and cancer research. These data may include thousands or millions of measurements on each item in the sample. For example, genotyping services provide individuals with information on hundreds of thousands of genetic variants in their cells and retailer databases may have information on the sales of tens of thousands of items for each store in the chain. Many of these data come in the form of counts (such as number of items of each type in inventory, number of mRNA molecules encoding a particular protein) or in the form of categories (such as on/off, present/absent, or genotype AA, aa or Aa). Methodology for highly multivariate continuous measurements such as blood pressure and temperature are well-developed but do not apply directly to count and categorical data. The investigator will develop statistical methodology and software to improve analysis and summary of count and categorical data. Four main areas of research are proposed: A) statistical models and tests to determine if the variables are associated with differences among groups; B) statistical methods for prediction or classification of group membership; C) methods to summarize the data with a much smaller set of derived variables which preserve the predictive power of the full data and D) multiple comparisons methods to estimate the error rates. For example, in a study of the genes associated with metastatic versus non-metastatic cancer, the methods could be used to determine which genes express differently in tumors which did or did not advance to metastasis, select a smaller set of genes which could be used as a diagnostic tool and then provide convenient summaries which can readily be interpreted by clinicians. In a study of stresses on a machine part, the pixels of scans of the part before and during the application of the stresses could be used to determine precise locations at which the part might fail and differences among features of the scan between parts which fail at low versus high stress. In studies in which a large number of models are fitted or tests conducted, it is necessary to tolerate a small percentage of errors. Concepts and methods in multiple testing which have been developed for continuous data will be extended to assist in estimating and controlling the number of false conclusions with count and categorical data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Mathematical Sciences Computing Research Environments
-
批准号:9627207
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:1996
-
负责人:Naomi Altman
-
依托单位:
Mathematical Sciences: Semi-parametric Methods for Longitudinal Data Analysis
-
批准号:9625350
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:1996
-
负责人:Naomi Altman
-
依托单位:
Mathematical Sciences: Computationally Intensive Problems in Statistics
-
批准号:8916245
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:1990
-
负责人:Naomi Altman
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: