课题基金 / 基金详情

CAREER: High-Dimensional Statistical Models for Unsupervised Learning

CAREER: High-Dimensional Statistical Models for Unsupervised Learning
职业:无监督学习的高维统计模型
批准号:
1945667
负责人:
Arash Amini
金额:
$40.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-07-15 至 2025-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
人们对数据和数据分析重要性的认识日益增强,再加上近年来数据量的空前增长,促使现在统称为数据科学领域的研究人员共同努力,开发能够处理大型复杂数据集的新模型。绝大多数可用数据都是未标记的,这使得建模问题更具挑战性。该项目将推动大型复杂未标记数据建模领域的发展。重点将放在从网络数据中学习以及从常规数据中学习依赖结构上。研究的一些具体问题是:通过网络传播的流行病可以告诉我们关于网络结构和流行病起源的什么信息?网络的结构能告诉我们哪些节点的潜在特征,例如,它们的分组,或者在社交网络的情况下的社区?在真实的网络中,除了简单的分组或社区结构之外,是否存在更精细的结构?我们能否从常规数据中学习复杂网络,这些数据告诉我们底层变量之间依赖关系的本质(例如,哪些变量是给定变量的原因)?这些复杂的模型对真实数据的拟合程度如何?在这些问题上取得进展对处理数据的许多科学领域具有直接影响。例如,基因组学和计算生物学、神经科学、流行病学、网络安全、社会科学和市场营销,都受益于网络分析的进步。依赖结构学习的进步可以改善因果推理程序,对所有科学领域都有影响。这个关于网络流行病的项目有可能立即应用于公共卫生领域,具有变革性。该项目推进了以无监督方式从数据中推断复杂关系的最新技术。因此,网络推理和图形建模将在我们的方法中发挥重要作用。我们将考虑四个主要任务:1)为结构化网络模型开发拟合优度测试,特别是那些用于社区检测和聚类的测试。尽管在网络建模方面取得了进步,但仍有人担心当前的模型无法捕捉真实网络的复杂性。实现现实网络建模的第一步是开发评估模型拟合程度的工具。2)推进了复杂网络建模技术的发展,提出了在真实网络中捕获自相似性的思路,以及多层网络的分层统计模型。3)基于网络动态的推进推理:许多网络都伴随着受网络结构支配的动态,如谣言、疾病的传播。我们经常观察动态的结果(谁会随着时间的推移而受到感染),并希望对动态的起源或底层网络的结构做出推断。我们将解决在现实网络中处理这些问题的挑战,其中许多周期的存在和关于动态的不完整信息构成了严重的困难。4)推进高维依赖结构的推断:描述随机变量集合之间的依赖关系(相关性、因果关系等)是统计分析的基本任务。首席研究员将探索从适合因果解释的数据中学习高维有向图形模型。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The growing awareness of the importance of data and data analysis, coupled with the unprecedented growth in the amount of data in recent years, has led to concerted efforts by researchers in the fields now collectively referred to as data sciences, to develop new models capable of handling big complex datasets. The vast majority of the available data is unlabeled, which makes the modeling problem more challenging. This project will advance the field of modeling big complex unlabeled data. The focus will be on learning from network data as well as learning dependency structures from regular data. Some of the concrete problems investigated are: What can an epidemic spreading over a network tell us about the structure of the network and the origin of the epidemic? What can the structure of the network tell us about the latent features of the nodes, for example, their grouping, or community in the case of social networks? Are there more refined structures in real networks beyond simple grouping or community structure? Can we learn complex networks from regular data that tell us about the nature of the dependency among the underlying variables (for example, what variables are the causes of a given variable)? How well do these often complex models fit the real data? Advancing on these questions has a direct impact on many scientific domains dealing with data. For example, genomics and computational biology, neuroscience, epidemiology, network security, social sciences and marketing, all benefit from advances in network analysis. Advances in dependency structure learning can improve causal inference procedures with impact on all scientific fields. This project on network epidemics has the potential to be transformative with immediate applications to the public health domain.This project advances the state-of-the art in inferring complex relations from data in an unsupervised fashion. As a result, network inference and graphical modeling will play prominent roles in our approach. We will consider four main tasks: 1) Developing goodness-of-fit tests for structured network models, in particular those used in community detection and clustering. Despite advances in network modeling, there are concerns that current models are not capturing the complexity of real networks. A first step toward realistic network modeling is developing tools for assessing how well the models fit. 2) Advancing the state-of-the-art in modeling complex networks, presenting ideas on capturing self-similarity in real networks as well as hierarchical statistical models for multilayer networks. 3) Advancing inference based on network dynamics: Many networks are accompanied by dynamics governed by the network structure, e.g., the spread of rumors and diseases. We often observe the result of the dynamics (who gets infected over time) and would like to make inference about the origin of the dynamic or the structure of the underlying network. We will address the challenges in dealing with these questions in real networks where the presence of many cycles and incomplete information about the dynamic pose serious difficulties. 4) Advancing inference of high-dimensional dependency structures: Characterizing dependencies (correlation, causation, etc.) among a collection of random variables is a fundamental task of statistical analysis. The principal investigator will explore learning high-dimensional directed graphical models from data that are suitable for causal interpretations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
On perfectness in Gaussian graphical models
论高斯图模型的完美性
DOI: --
发表时间: 2022
期刊: Proceedings of The 25th International Conference on Artificial Intelligence and Statistics
影响因子: --
作者: [Amini, Arash A., Aragam, Bryon, Zhou, Qing]
通讯作者: Zhou, Qing
DOI: 10.1214/22-ba1355
发表时间: 2019-03
期刊: ArXiv
影响因子: --
作者: [M. Paez;A. Amini;Lizhen Lin]
通讯作者: M. Paez;A. Amini;Lizhen Lin
DOI: --
发表时间: 2021
期刊:
影响因子: --
作者: [Linfan Zhang;A. Amini]
通讯作者: Linfan Zhang;A. Amini
DOI: --
发表时间: 2020
期刊:
影响因子: --
作者: [Zahra S. Razaee;A. Amini]
通讯作者: Zahra S. Razaee;A. Amini
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis