课题基金 / 基金详情

CAREER: Learning Probabilistic Factor Models

CAREER: Learning Probabilistic Factor Models
职业:学习概率因子模型
批准号:
1943902
负责人:
Zheng Ke
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-07-01 至 2025-06-30

项目摘要

项目成果

Zheng Ke的其他基金

相似基金

相关文献

中文摘要
翻译
大量的文本和社交网络数据出现在科学研究和日常生活中。该项目将开发用于分析数据的统计方法,从而产生新的科学、社会学和生物医学发现。由于数据的特点,该研究面临着几个根本性的挑战:(1)大规模,这需要先进的存储、计算和质量控制策略;(2)结构复杂,需要进行细致的统计建模;(3)强噪声,这需要复杂的降噪技术。为了应对这些挑战,PI提出了一种通用的概率因子建模方法。这项研究将为社会网络分析、自然语言处理、rna测序数据分析和电子健康记录分析提供一系列统计工具。该项目还将帮助培训研究生和本科生的数据收集、数据清理、统计方法和理论。此外,该项目将发布新的网络和文本分析软件和数据集,为教育和研究提供有用的资源。概率因子模型是指因子或因子载荷与概率质量函数相关联的因子模型。例子包括文本挖掘中的主题模型和社交网络中的混合成员模型。由于这些模型中的非负约束和相关异方差噪声,使得统计估计和推断极具挑战性。本项目将解决这些挑战,并将提出的方法应用于不同的应用。第一个目标是开发一个新的框架来探索主题模型中的稀疏性。它提出了一种新的词汇“稀疏性”概念,不同于高维统计中传统的稀疏性概念。该框架将为文本挖掘中的降维提供理论基础,并为主题权重估计提供新的词筛选方法和新的谱方法。第二部分研究网络混合隶属度估计的基本统计限制。这将产生一种新的混合隶属度估计的最优性理论,特别是对于具有很大程度异质性的网络模型,以及新的经验特征向量随机矩阵理论。它还将产生有关统计相关领域学术研究人员之间网络的数据集,并产生关于学术研究趋势和模式的发现。第三个重点是使上述技术工具适应生物医学数据,包括大量和单细胞rna测序数据和电子卫生保健数据。它将产生新的混合模型和生物医学数据的统计推断工具。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
A large amount of text and social network data is emerging in scientific research as well as everyday life. This project will develop statistical methods for analyzing data resulting in new scientific, sociological, and biomedical discoveries. The research has several fundamental challenges due to the features of the data: (1) large scale, which requires advanced strategies on storage, computation, and quality control; (2) a complicated structure, which makes careful statistical modeling a critical need; and (3) strong noise, which requires sophisticated de-noising techniques. To address these challenges, the PI proposes a universal probabilistic factor modeling approach. The research will provide an array of statistical tools for social network analysis, natural language processing, RNA-sequencing data analysis, and electronic health records analysis. This project will also help train graduate and undergraduate students on data collection, data cleaning, statistical methodology and theory. In addition, this project will release new software and data sets for network and text analysis providing useful resources for both education and research. Probabilistic factor models refer to factor models whose factors or factor loadings are connected to probability mass functions. Examples include the topic models in text mining and mixed membership models in social networks. Due to the nonnegative constraints and the dependent and heteroscedastic noise in these models, statistical estimation and inference are extremely challenging. This project will tackle these challenges and apply the proposed methods to different applications. The first thrust aims to develop a novel framework for exploring sparsity in topic models. It proposes a new notion of "sparsity" on the vocabulary, which is different from the conventional notion of sparsity in high-dimensional statistics. The framework will provide a theoretical foundation for dimension reduction in text mining, as well as new word screening methods and new spectral methods for topic weight estimation. The second thrust aims to study the fundamental statistical limits for network mixed membership estimation. It will lead to a new optimality theory of mixed membership estimation, especially for network models with a large degree of heterogeneity, and new random matrix theory for empirical eigenvectors. It will also produce data sets about the networks among academic researchers in statistics-related fields and generate discoveries about the trend and patterns in academic research. The third thrust aims to adapt the above technical tools to biomedical data, including bulk and single-cell RNA-sequencing data and electronic health care data. It will result in new mixture models and statistical inference tools for biomedical data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
Improvements on SCORE, Especially for Weak Signals
SCORE 的改进,特别是对于弱信号
DOI: 10.1007/s13171-020-00240-1
发表时间: 2021
期刊: Sankhya A
影响因子: --
作者: [Jin, Jiashun, Ke, Zheng Tracy, Luo, Shengming]
通讯作者: Luo, Shengming
Phase transition for detecting a small community in a large network
用于检测大型网络中的小社区的相变
DOI: --
发表时间: 2023
期刊: The Eleventh International Conference on Learning and Representations
影响因子: --
作者: [Jin, Jiashun, Ke, Zheng Tracy, Turner, Paxton, Zhang, Anru]
通讯作者: Zhang, Anru
DOI: 10.1214/21-aos2089
发表时间: 2019-04
期刊: The Annals of Statistics
影响因子: --
作者: [Jiashun Jin;Z. Ke;Shengming Luo]
通讯作者: Jiashun Jin;Z. Ke;Shengming Luo
A Comparison of Hamming Errors of Representative Variable Selection Methods
代表性变量选择方法的汉明误差比较
DOI: --
发表时间: 2022
期刊: ICLR 2022
影响因子: --
作者: [Ke, Zheng Tracy, Wang, Longlin]
通讯作者: Wang, Longlin
共 13 条
    Hidden Components in Modern Applications
    • 批准号:
      1925845
    • 项目类别:
      Standard Grant
    • 资助金额:
      $14.18万
    • 财政年份:
      2018
    • 负责人:
      Zheng Ke
    • 依托单位:
    Hidden Components in Modern Applications
    • 批准号:
      1712958
    • 项目类别:
      Standard Grant
    • 资助金额:
      $20.0万
    • 财政年份:
      2017
    • 负责人:
      Zheng Ke
    • 依托单位:
    国内基金
    海外基金
    Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
    Understanding structural evolution of galaxies with machine learning
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      Nicola Rosario Napolitano
    • 依托单位:
    煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
    • 批准号:
      --
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      30万元
    • 批准年份:
      2022
    • 负责人:
      吉建娇
    • 依托单位:
    基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
    • 批准号:
      62003314
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      24.0万元
    • 批准年份:
      2020
    • 负责人:
      沈剑
    • 依托单位: