Robust and scalable algorithms for learning hidden structures in sparse network data with the aid of side information
Robust and scalable algorithms for learning hidden structures in sparse network data with the aid of side information
批准号:
2311024
负责人:
Adel Javanmard
金额:
$27.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-15 至 2026-07-31
中文摘要
网络数据在各个领域变得越来越重要,包括社会科学、生物学、计算机科学和工程学。在社会科学中,网络数据被用来研究社会互动和关系,如友谊网络、政治关系和知识转移网络。在生物学中,网络数据用于建模和分析生物系统,如基因调控网络、蛋白质-蛋白质相互作用网络和食物网。学习网络中的隐藏结构,特别是社区结构的检测和建模,是至关重要的。这一过程不仅增强了数据的可解释性,而且还实现了数据压缩,通过检测潜在的子种群并为每个子种群拟合合适的模型来管理数据异质性,并解决了缺失标签的问题。尽管有大量的聚类算法,但当前的方法往往存在可伸缩性和健壮性问题,限制了它们在实际应用程序中的有效性。此外,随着数据共享变得越来越普遍,通常存在大量关于网络中节点的上下文信息,例如在线平台上用户的人口统计信息或浏览历史记录,这些信息可以有效地与网络数据相结合,从而大大提高聚类过程的有效性。为了解决这些限制并推动该领域的发展,该项目将开发健壮且可扩展的推理网络方法,该方法适应节点度的异质性,并允许将节点侧信息与成对交互数据相结合,以进行更有效的分析。该项目将通过为来自不同背景的研究生提供参与前沿研究的培训机会,支持统计和机器学习研究方面的教育。它还通过提供理解和管理复杂网络系统的工具而造福社会。本研究由三个相互关联的部分组成,这三个部分协同工作,为稀疏网络数据中潜在结构的学习提供了一个统一的框架。在第一部分中,我们设计了基于半确定规划的聚类算法,该算法允许我们将节点上的高维上下文信息与交互图结合起来。在第二部分中,我们将开发方法来提高推理算法对节点上下文数据或交互图中的对抗性扰动的鲁棒性。第三部分以前两部分设计的算法为基础,提供了这些算法的低计算和内存效率实现,可以扩展到大规模网络数据。此外,该项目将调查该项目在不同领域的潜在用途,利用所得的聚类算法和优化统计工具。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Network data is becoming increasingly relevant in various areas, including social sciences, biology, computer science, and engineering. In the social sciences, network data is used to study social interactions and relationships, such as friendship networks, political affiliations, and knowledge transfer networks. In biology, network data is used to model and analyze biological systems, such as gene regulation networks, protein-protein interaction networks, and food webs. Learning the hidden structures within networks, in particular detecting and modeling community structures, is of paramount importance. This process not only enhances the interpretability of data but also enables data compression, manages data heterogeneity by detecting latent subpopulations and fitting appropriate models to each, and addresses the issue of missing labels. Despite a plethora of clustering algorithms, current approaches often suffer from scalability and robustness issues, limiting their effectiveness in real-world applications. Furthermore, as data sharing becomes more prevalent, there is often a wealth of contextual information available about the nodes in a network, such as demographic information or browsing history for users on an online platform, that can be effectively combined with network data to greatly enhance the effectiveness of clustering procedures. To tackle these limitations and advance the field, this project will develop robust and scalable inferential network methods, which adapt to the heterogeneity of node degrees and allows to combine nodewise side information with pairwise interaction data for a more effective analysis. This project will support education in statistical and machine learning research by providing training opportunities for graduate students, from diverse backgrounds, to participate in cutting-edge research. It also benefits society by providing tools for understanding and managing complex network systems.This research consists of three interrelated parts, which work in concert to provide a unifying framework for learning latent structures in sparse network data. In the first part, we devise clustering algorithms based on semidefinite programming which allows us to combine the high-dimensional contextual information on the nodes with the interaction graph. In the second part, we will develop methods to improve the robustness of the inferential algorithms to adversarial perturbations in the nodewise contextual data or the interaction graph. The third part builds on the algorithms devised in the previous two parts and provides low-computation and memory-efficient implementation of these algorithms that can scale to large-scale network data. In addition, the project will investigate the potential uses of this project across diverse domains, utilizing the resulting clustering algorithms and optimization-statistical tools.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Valid and Scalable Inference for High-dimensional Statistical Models
-
批准号:1844481
-
项目类别:Continuing Grant
-
资助金额:$40.22万
-
财政年份:2019
-
负责人:Adel Javanmard
-
依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位: