课题基金 / 基金详情

CAREER: Methodology for Statistical Computing in Massive Datasets: Parallel Approaches to Cluster and MCMC Estimation

CAREER: Methodology for Statistical Computing in Massive Datasets: Parallel Approaches to Cluster and MCMC Estimation
职业:海量数据集中的统计计算方法:聚类和 MCMC 估计的并行方法
批准号:
0437555
负责人:
Ranjan Maitra
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-07-01 至 2010-11-30

项目摘要

项目成果

Ranjan Maitra的其他基金

相似基金

相关文献

中文摘要
翻译
职业:大规模数据集统计计算方法:聚类和MCMC估计的并行方法DMS 0239734 PI:Ranjan Maitra该项目旨在开发大规模数据库统计分析和估计的实用方法。 由于自动化的数据收集方法,现在有一个严重的多维记录过剩。 将它们分成同质簇以更好地理解它们是各种应用中的理想目标,但经典的统计方法在许多情况下在计算上是不可行的。 我建议为这方面制定平行的方法。 虽然开发的方法和理论将是相当普遍的,并在商业,医学,环境,软件质量评估等应用程序中的潜在用途,我将在三个科学合作的背景下进行研究。 第一个涉及美国环境保护署(EPA)自我报告的有毒物质释放清单(TRI)数据库,在这些数据库中,根据其产品,人口统计和商业信息对不同设施进行分析可以提高记录的准确性,并更好地描述它们的排放组合。 第二个项目是评估功能性磁共振成像(fMRI)扫描的可靠性,以了解大脑的认知过程,作为患者护理和治疗的第一步。 第三个应用是生物信息学,其目标是聚类微阵列数据,并分析二维蛋白质组凝胶图像。 这将有助于分离基因并了解它们与不同疾病的关系。 聚类通常是一个非常困难的问题,即使对于非常中等大小的数据集也有经验解决方案。 我建议在几种不同的情况下开发多通道方法。 我还提出了多尺度模拟方法来估计在高维背景下。 由于高维性,仿真方法所面临的最大挑战之一是由于其广阔而要遍历的空间周围的低移动性。 我建议通过将这些高维空间连接到低维空间(低维空间要小得多),并使用这些低尺度从高维空间的一个角落遍历到另一个角落来解决这个问题。 这个五年计划的最终目标是研究蛋白质组凝胶数据更复杂模型的开发和估计。 提出的大多数计划只有通过并行计算接口才有可能实现。 这在大量的科学应用中越来越重要,我建议通过设计合适的量身定制的研究生和本科生课程,同时为统计学生提供必要的专业知识。
英文摘要
CAREER: Methodology for Statistical Computing in Massive Datasets: Parallel Approaches to Cluster and MCMC EstimationDMS 0239734PI: Ranjan MaitraThis project is aimed at developing practical methodology for statistical analysis and estimation in massively sized databases. Because of automated data collection methods, there is nowa surfeit of severely multi-dimensional records. Grouping them into homogeneous clusters to better understand them is a desirable goal in a variety of applications, yet classical statistical methods are computationally infeasible in many cases. I propose to develop parallel methodology for this context. Although the methodology and theory developed will be quite general and for potential use in applications ranging from business, medicine, the environment, software quality assessment, I will conduct the research in the context of three scientific collaborations. The first pertains to theUS Environmental Protection Agency's (EPA) self-reported Toxic Releases Inventory (TRI) databases, where profiling the different facilities in terms of their product, demographic and business information can improve the accuracy of records, as well as better characterize them vis-a-vis their emissions mix. The second project is to assess the reliability of functional Magnetic Resonance Imaging (fMRI) scans, with a view to understanding the cognitive processes of the brain, as a first step to patient care and therapy. The third application is in bioinformatics where the goal is to cluster microarray data and also to analyze two-dimensional proteomic gel images. This will help in isolating genes and understanding their relationship with different disorders. Clustering is in general a very difficult problem, with empirical solutions even for very moderately sized datasets. I propose to develop multi-pass methodologies in several different scenarios. I also propose multi-scale simulation approaches to estimation in a high-dimensional context. One of the biggest challenges faced by simulation methods due to high dimensionality is the low mobility around the space to be traversed because of its vastness. I propose to address this issue by connecting these high-dimensional spaces to lower-dimensional ones (which are significantly smaller) and by using these lower scales to traverse from one corner of the higher-dimensional space to another. A final goal of this five-year plan is to investigate the development and estimation in more complex models for proteomic gel data. Most of the plans proposed will be possible only with a parallel computing interface. This is increasingly critical in a large number of scientific applications, and I propose to simultaneously provide statistics students with the necessary expertise by designing suitably tailored graduate and undergraduate classes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Methodology for Statistical Computing in Massive Datasets: Parallel Approaches to Cluster and MCMC Estimation
海外基金