课题基金 / 基金详情

BIGDATA: Collaborative Research: F: Big Data, It's Not So Big: Exploiting Low-Dimensional Geometry for Learning and Inference

BIGDATA: Collaborative Research: F: Big Data, It's Not So Big: Exploiting Low-Dimensional Geometry for Learning and Inference
BIGDATA:协作研究:F:大数据,它并不是那么大:利用低维几何进行学习和推理
批准号:
1546132
负责人:
Shayn Mukherjee
金额:
$32.22万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-12-01 至 2018-11-30

项目摘要

项目成果

Shayn Mukherjee的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This research will leverage ideas from algebraic and differential geometry to address core problems in modern high-dimensional and massive data science. The project will develop statistical methods and numerical tools, grounded in solid mathematical, statistical, and computational foundations, to extract low dimensional geometry from massive data with applications in clustering, data summarization, prediction, dimension reduction, and visualization. The solutions developed as part of this project can result in fundamental advances in practical applications across fields as diverse as biology, medicine, social sciences, communication networks, and engineering. In addition to internal validation via statistical and mathematical theory and simulation studies, the methods developed in the project will involve external validation via interdisciplinary applications. These applications include: (1) inference of population structure from genomic data; (2) document analysis via topic models; and (3) inference of subsets of putative gene networks relevant to drug resistance in melanoma.The research is motivated by the central premise that, even though the amount of data may be massive, a compact model can represent these data. Specifically, high-dimensional and/or massive data can be reasonably approximated by a mixture of subspaces, for which sparse representations exist. A mixture of subspaces of potentially different dimensions is a flexible, rich representation of data with nice mathematical properties that can scale to large data. There are several fundamental challenges in modeling mixtures of subspaces that will be addressed in this research: 1) the subspaces will be of different dimensions, 2) both the subspace parameters and the mixing parameters need to be inferred, 3) efficient algorithms for inference are required for both high-dimensional and massive data. The central foundational impediment in all of these challenges is that the model is a stratified space (a union of manifolds), and therefore has singularities. The key insight in this research is that there exist embeddings and representations of the model space that mitigate these singularities. These ideas are implemented as concrete Bayesian, frequentist, and numerical algorithms and models to address the real world examples listed above.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
HDR TRIPODS: Innovations in Data Science: Integrating Stochastic Modeling, Data Representations, and Algorithms
  • 批准号:
    1934964
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $150.0万
  • 财政年份:
    2019
  • 负责人:
    Shayn Mukherjee
  • 依托单位:
Beyond Riemannian Geometry in Inference
  • 批准号:
    1713012
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $22.0万
  • 财政年份:
    2017
  • 负责人:
    Shayn Mukherjee
  • 依托单位:
Collaborative Research: Topological Methods for Parsing Shapes and Networks and Modeling Variation in Structure and Function
  • 批准号:
    1418261
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $31.12万
  • 财政年份:
    2014
  • 负责人:
    Shayn Mukherjee
  • 依托单位:
Collaborative Research: Numerical algebra and statistical inference
  • 批准号:
    1209155
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2012
  • 负责人:
    Shayn Mukherjee
  • 依托单位:
海外基金