课题基金 / 基金详情

III: Small-Collaborative: Efficient Bayesian Model Computation for Large and High Dimensional Data Sets

III: Small-Collaborative: Efficient Bayesian Model Computation for Large and High Dimensional Data Sets
III:小型协作:大型高维数据集的高效贝叶斯模型计算
批准号:
0914861
负责人:
Carlos Ordonez
金额:
$33.9万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-08-01 至 2013-07-31

项目摘要

项目成果

Carlos Ordonez的其他基金

相似基金

相关文献

中文摘要
翻译
该奖项是根据2009年美国复苏和再投资法案(公法111-5)资助的。该资助支持适应和优化马尔可夫链蒙特卡罗方法的研究,以计算驻留在二级存储上的大型数据集的贝叶斯模型,利用数据库系统技术。这项工作将寻求优化计算,保持模型准确性,加速从大型和高维数据集采样技术,利用不同的数据集布局和索引数据结构。该团队将开发加权抽样方法,这种方法可以产生与传统抽样方法质量相似的模型,但对于不能放在主存储上的大型数据集来说,这种方法要快得多。一个子目标将研究如何压缩一个大数据集,以保持参数贝叶斯模型的统计属性,然后适应现有的方法来处理压缩数据集。智力价值和更广泛的影响这一努力需要开发新的计算方法,可以有效地处理大数据集和数字密集型计算。主要的技术困难是不可能从大数据集的子样本中获得准确的样本。因此,团队将专注于基于整个数据集的后验分布加速采样。这个问题异常困难,因为随机方法需要在整个数据集上进行大量的迭代(通常是数千次)才能收敛。然而,如果数据集被压缩,就有必要将传统的方法推广到使用加权点与高阶统计量相结合,而不是众所周知的高斯分布的充分统计量。开发结合主存储和辅助存储的优化与优化仅在主存储上工作的算法有很大不同。这项研究工作需要贝叶斯模型和随机方法的综合统计知识,超越传统的数据挖掘方法。在优化使用大型磁盘驻留矩阵的计算方面也需要强大的数据库系统背景。与现代统计软件包解决随机模型相比,这项研究将能够更快地解决更大规模的问题。贝叶斯分析和模型管理将更容易、更快、更灵活。广泛影响本研究将在三个独立的应用领域进行:癌症、水污染和癌症和心脏病患者的医疗数据集。这项拨款的教育部分将加强当前数据挖掘的教学和研究。在高级数据挖掘课程中,学生将应用随机方法计算数百个变量和数百万条记录的复杂贝叶斯模型。数据挖掘研究项目将加强贝叶斯模型,促进统计学与计算机科学之间的互动。关键词:贝叶斯模型,随机方法,数据库系统
英文摘要
This award is funded under the American Recovery and Reinvestment Act of 2009 (Public Law 111-5.This grant supports research in adapting and optimizing Markov Chain Monte Carlo methods to compute Bayesian models on large data sets resident on secondary storage, exploiting database systems techniques. The work will seek to optimize computations, preserve model accuracy and accelerate sampling techniques from large and high dimensional data sets, exploiting different data set layouts and indexing data structures. The team will develop weighted sampling methods that can produce models of similar quality as traditional sampling methods, but which are much faster for large data sets that cannot fit on primary storage. One sub-goal will study how to compress a large data set preserving its statistical properties for parametric Bayesian models, and then adapting existing methods to handle compressed data sets. Intellectual Merit and Broader ImpactThis endeavor requires developing novel computational methods that can work efficiently with large data sets and numerically intensive computations. The main technical difficulty is that it is not possible to obtain accurate samples from subsamples of a large data set. Therefore, the team will focus on accelerating sampling from the posterior distribution based on the entire data set. This problem is unusually difficult because stochastic methods require a high number of iterations (typically thousands) over the entire data set to converge. However, if the data set is compressed it becomes necessary to generalize traditional methods to use weighted points combined with higher order statistics, beyond the well-known sufficient statistics for the Gaussian distribution. Developing optimizations combining primary and secondary storage is quite different from optimizing an algorithm that works only on primary storage. This research effort requires comprehensive statistical knowledge on both Bayesian models and stochastic methods, beyond traditional data mining methods. A strong database systems background in optimizing computations with large disk-resident matrices is also necessary. This research will enable a faster solution of larger scale problems compared to modern statistical packages to solve stochastic models. Bayesian analysis and model management will be easier, faster and more flexible. Broad ImpactThis research will occur within the context of three separate application areas: cancer, water pollution, and medical data sets with patients having cancer and heart disease. The educational component of this grant will enhance current teaching and research on data mining. In an advanced data mining course students will apply stochastic methods to compute complex Bayesian models on hundreds of variables and millions of records. Data mining research projects will be enhanced with Bayesian models, promoting interaction between statistics and computer science.Keywords: Bayesian model, stochastic method, database system
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Equilibria of Two Relaxed Plasma Species With One Species Confined by the Space Charge of the Other Species
  • 批准号:
    1803047
  • 项目类别:
    Standard Grant
  • 资助金额:
    $18.6万
  • 财政年份:
    2018
  • 负责人:
    Carlos Ordonez
  • 依托单位:
Collaborative Research: Experimental and Theoretical Study of the Plasma Physics of Antihydrogen Generation and Trapping
  • 批准号:
    1500427
  • 项目类别:
    Standard Grant
  • 资助金额:
    $22.8万
  • 财政年份:
    2015
  • 负责人:
    Carlos Ordonez
  • 依托单位:
Collaborative Research: Experimental and Theoretical Study of the Plasma Physics of Antihydrogen Generation and Trapping
  • 批准号:
    1202428
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2012
  • 负责人:
    Carlos Ordonez
  • 依托单位:
Pan-American Institute of Science and Technology
  • 批准号:
    0936560
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.0万
  • 财政年份:
    2009
  • 负责人:
    Carlos Ordonez
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: