Statistical Methods for High Dimensional Discrete Data
Statistical Methods for High Dimensional Discrete Data
批准号:
1007801
负责人:
Naomi Altman
金额:
$20.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-06-01 至 2015-05-31
中文摘要
现在,非常高的维数和二进制数据在许多领域都很常见,包括机器学习、成像和营销。 在高通量生物学中,产生计数和分类数据的超高通量测序技术正在取代微阵列和其他“组学”测量设备。这些测量设备的输出是针对每个样本数万个响应的每个基因或其他生物亚基的计数,或者针对每个样本可能数百万个响应的单核苷酸多态性 (SNP) 等特征的存在/不存在。 类似的数据可以从卫星图像、医学扫描、监控设备和其他高维测量设备的特征中得出。 研究人员会将针对连续(主要正态分布)数据开发的高度多元和多重测试方法扩展到离散数据。 新方法将在四个领域开发:A)分析离散数据的分布差异,这些数据可以使用具有过度分散和贝叶斯或经验贝叶斯收缩的广义线性混合模型来适应复杂的实验设计。 B) 考虑离散预测变量的误差结构,对离散数据设置中的样本和变量进行监督聚类的方法。 C)经典且足够的降维方法,例如离散数据的典型相关和切片逆回归。 D) 多重测试中概念和方法的扩展,例如对离散设置的错误发现率估计,其中使用条件混合建模来自独立或弱相关测试的 p 值可能具有不同的零分布。 这些方法将在基因组学和成像数据上进行测试。高度多元的数据现在已成为细胞生物学、营销、医学和卫星成像、气象学、流行病学、欺诈检测和癌症研究等多个领域的常态。 这些数据可能包括样本中每个项目的数千或数百万次测量。 例如,基因分型服务为个人提供有关其细胞中数十万种遗传变异的信息,零售商数据库可能包含连锁店中每家商店数万种商品的销售信息。许多这些数据以计数的形式出现(例如库存中每种类型的项目数量、编码特定蛋白质的 mRNA 分子的数量)或以类别的形式出现(例如开/关、存在/不存在或基因型 AA、aa 或 Aa)。 血压和体温等高度多元连续测量的方法已经很成熟,但并不直接应用于计数和分类数据。 研究者将开发统计方法和软件,以改进计数和分类数据的分析和总结。 提出了四个主要研究领域:A)统计模型和测试,以确定变量是否与组间差异相关; B) 用于预测或分类群体成员资格的统计方法; C) 用更小的派生变量集总结数据的方法,保留完整数据的预测能力;D) 估计错误率的多重比较方法。 例如,在与转移性癌症和非转移性癌症相关的基因的研究中,这些方法可用于确定哪些基因在进展或未进展为转移的肿瘤中表达不同,选择可用作诊断工具的较小基因组,然后提供临床医生可以轻松解释的方便的总结。 在机器零件应力的研究中,在施加应力之前和期间零件的扫描像素可用于确定零件可能失效的精确位置以及在低应力与高应力下失效的零件之间的扫描特征之间的差异。 在拟合大量模型或进行测试的研究中,有必要容忍一小部分错误。 针对连续数据开发的多重测试的概念和方法将得到扩展,以帮助估计和控制计数和分类数据的错误结论的数量。
英文摘要
Very high dimensional count and binary data are now common in many fields including machine learning, imaging and marketing. In high-throughput biology, ultra-high thoughput sequencing technologies which produce count and categorical data are displacing microarrays and other "omics" measurement devices. The output of these measurement devices are counts per gene or other biological subunit for tens of thousands of responses per sample, or presence/absence for features such as single nucleotide polymorphisms (SNPs), for possibly millions of responses per sample. Similar data can be derived on for features on satellite images, medical scans, monitoring devices and other very high dimensional measurement devices. The investigator will extend highly multivariate and multiple testing methods developed for continuous (primarily normally distributed) data to discrete data. New methods will be developed in four areas: A) analyses for differences in distribution for discrete data that can accommodate complex experimental designs using generalized linear mixed models with overdispersion and Bayesian or empirical Bayes shrinkage. B) methods for supervised clustering of samples and variables in the discrete data setting taking into account the error structure of the discrete predictors. C) classical and sufficient dimensions reduction methods such as canonical correlation and sliced inverse regression for discrete data. D) extension of concepts and methods in multiple testing, such as false discovery rate estimation to the discrete setting in which the p-values from independent or weakly dependent tests may have different null distributions using conditional mixture modeling. The methods will be tested on genomics and imaging data.Very highly multivariate data are now the norm in fields as diverse as cell biology, marketing, medical and satellite imaging, meteorology, epidemiology, fraud detection and cancer research. These data may include thousands or millions of measurements on each item in the sample. For example, genotyping services provide individuals with information on hundreds of thousands of genetic variants in their cells and retailer databases may have information on the sales of tens of thousands of items for each store in the chain. Many of these data come in the form of counts (such as number of items of each type in inventory, number of mRNA molecules encoding a particular protein) or in the form of categories (such as on/off, present/absent, or genotype AA, aa or Aa). Methodology for highly multivariate continuous measurements such as blood pressure and temperature are well-developed but do not apply directly to count and categorical data. The investigator will develop statistical methodology and software to improve analysis and summary of count and categorical data. Four main areas of research are proposed: A) statistical models and tests to determine if the variables are associated with differences among groups; B) statistical methods for prediction or classification of group membership; C) methods to summarize the data with a much smaller set of derived variables which preserve the predictive power of the full data and D) multiple comparisons methods to estimate the error rates. For example, in a study of the genes associated with metastatic versus non-metastatic cancer, the methods could be used to determine which genes express differently in tumors which did or did not advance to metastasis, select a smaller set of genes which could be used as a diagnostic tool and then provide convenient summaries which can readily be interpreted by clinicians. In a study of stresses on a machine part, the pixels of scans of the part before and during the application of the stresses could be used to determine precise locations at which the part might fail and differences among features of the scan between parts which fail at low versus high stress. In studies in which a large number of models are fitted or tests conducted, it is necessary to tolerate a small percentage of errors. Concepts and methods in multiple testing which have been developed for continuous data will be extended to assist in estimating and controlling the number of false conclusions with count and categorical data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Mathematical Sciences Computing Research Environments
-
批准号:9627207
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:1996
-
负责人:Naomi Altman
-
依托单位:
Mathematical Sciences: Semi-parametric Methods for Longitudinal Data Analysis
-
批准号:9625350
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:1996
-
负责人:Naomi Altman
-
依托单位:
Mathematical Sciences: Computationally Intensive Problems in Statistics
-
批准号:8916245
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:1990
-
负责人:Naomi Altman
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: