课题基金 / 基金详情

EAGER: Adaptive Sampling of Massive Graph Streams

EAGER: Adaptive Sampling of Massive Graph Streams
EAGER:海量图流的自适应采样
批准号:
1848596
负责人:
Nicholas Duffield
金额:
$20.05万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2022-08-31

项目摘要

项目成果

Nicholas Duffield的其他基金

相似基金

相关文献

中文摘要
翻译
作为日常生活基础的电子服务-如互联网、社交网络、银行系统和在线零售商-不断产生大量关于其操作和使用的数据。服务提供商的一项关键分析任务是了解不同数据点之间的共性,例如根据社交互动模式提供朋友推荐,或者了解一组相关交易如何共同表示银行欺诈。产生这种数据的数量给数据的存储和处理带来了挑战。另一方面,采样和其他数据简化技术阻碍了识别取证和其他回顾性分析模式的能力。 虽然以前的工作已经提出了解决方案,在特定的情况下,目前很少有一般的理解如何优化选择的采样算法,以最好地匹配的要求下的数据分析的操作资源的限制。此外,虽然现有的解决方案已经调整了他们的专家设计师,在没有一个通用的框架,它是具有挑战性的研究人员在应用领域,以适应他们的问题。为了弥补这些差距,该项目建立了一个新的通用框架,用于流图数据的最佳采样。它将使用这个框架来实现问题规范和采样算法之间的映射,以简化应用程序工作代码的生成。该项目还将编写关于数据流采样的教育材料,并向应用领域的研究人员进行外联,以促进所开发成果的适用性和使用,这项工作将为图形流采样及其应用开发一个变革性框架,以解决知识和实践方面的突出差距。首先,虽然目前的方法经常集中在全局图属性的计算上,但应用程序通常需要一个代表性的样本进行快速或回顾性分析,而不必重新处理整个流,即使可用。其次,当前的方法通常针对特定的子图目标和问题进行优化,而没有针对分析精度和资源需求进行优化的一般能力。第三,它是具有挑战性的领域研究人员推广这些结果超出了原来的问题,因为他们的创作所需的专家调整。作为回应,该项目将开发一个框架图数据流采样,适用于广泛的应用程序问题和设置,允许可调的准确性和空间和时间计算资源之间的权衡,并可实现为问题规范和工作应用程序代码之间的映射。具体而言,这项工作将使用一种新形式的加权水库优先级采样,以适应采样和保留的边缘,以不断演变的作用,在采样图和分析目标。该框架将使时间衰减的聚类子图采样,和多重图采样,以支持从流的重复边缘估计。在一个有代表性的各种应用程序的分析和评估将被用来优化计算策略,允许灵活的准确性和资源消耗之间的权衡。这些结果将通过实现框架作为问题规范和工作应用程序代码之间的映射来改变抽样方法的使用,使领域研究人员能够快速将这项工作的结果纳入他们自己的应用程序中。该奖项反映了NSF的法定使命,并被认为值得通过使用基金会的智力价值和更广泛的影响审查标准进行评估来支持。
英文摘要
The electronic services that underpin daily life---such as the internet, social networks, banking systems, and online retailers---continuously generate vast amounts of data concerning their operation and usage. A key analytical task for service providers is to understand the commonality between different data points, e.g. to provide friend recommendations based on patterns of social interaction, or to understand how a set of related transactions may collectively signify a bank fraud. The volume at which such data is generated presents challenges for its storage and processing. On the other hand, sampling and other data reduction techniques hinder the ability to discern patterns for forensics and other retrospective analysis. While previous work has proposed solutions in specific instances, there is currently little general understanding of how to optimize the choice of sampling algorithm to best match the requirement of data analysis under operational resource constraints. Furthermore, while existing solutions have been tuned by their expert designers, in the absence of a general framework it is challenging for researchers in application domains to adapt these to their problems. To meet these gaps, this project establishes a new general framework for optimal sampling of streaming graph data. It will use this framework to implement a map between problem specifications and sampling algorithms in order to streamline the generation of working code for applications. The project will also develop educational materials on data stream sampling and outreach to researchers in application domains to promote applicability and usage of the results developed.This work will develop a transformative framework for graph stream sampling and its applications that address outstanding gaps in knowledge and practice. First, while current methods frequently focus on computation of global graph properties, applications often require a representative sample for rapid or retrospective analysis without having to reprocess the entire stream, even if available. Second, current methods are typically optimized for specific subgraph targets and problems without a general ability to optimize for analysis accuracy and resource requirements. Third, it is challenging for domain researchers to generalize these results beyond the original problem because of the expert tuning required for their creation. In response, this project will develop a framework graph data stream sampling that is applicable to a wide class of application problems and settings, allows tunable trade-off between accuracy and space and time computational resources, and is implementable as a mapping between problem specification and working application code. Specifically, the work will use a novel form of weighted reservoir priority sampling to adapt the sampling and retention of edges to their evolving role in the sampled graph and with respect to analysis goals. The framework will enable time-decayed clustered subgraph sampling, and multigraph sampling to support estimation from streams of repeated edges. Analysis and evaluation in a representative variety of applications will be used to optimize computational strategies allowing flexible trade-off between accuracy and resource consumption. These results will transform usage of sampling methods by realizing the framework as a mapping between problem specification and working application code that enables domain researchers to rapidly incorporate the results of this work within their own applications.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(12)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3442202
发表时间: 2021-04
期刊: ACM Transactions on Knowledge Discovery from Data (TKDD)
影响因子: --
作者: [Nesreen K. Ahmed;N. Duffield;Ryan A. Rossi]
通讯作者: Nesreen K. Ahmed;N. Duffield;Ryan A. Rossi
DOI: --
发表时间: 2019-08
期刊:
影响因子: --
作者: [Ehsan Hajiramezanali;Arman Hasanzadeh;N. Duffield;K. Narayanan;Mingyuan Zhou;Xiaoning Qian]
通讯作者: Ehsan Hajiramezanali;Arman Hasanzadeh;N. Duffield;K. Narayanan;Mingyuan Zhou;Xiaoning Qian
DOI: 10.1109/icdm.2018.00043
发表时间: 2018-08
期刊: 2018 IEEE International Conference on Data Mining (ICDM)
影响因子: --
作者: [Xi Liu;Muhe Xie;Xidao Wen;Rui Chen;Yong Ge;N. Duffield;Na Wang]
通讯作者: Xi Liu;Muhe Xie;Xidao Wen;Rui Chen;Yong Ge;N. Duffield;Na Wang
DOI: --
发表时间: 2020
期刊: Advances in neural information processing systems
影响因子: --
作者: [Ahmed, Nesreen, Duffield, Nick]
通讯作者: Duffield, Nick
共 10 条
    EAGER: Real-Time: Learning-Mediated Control for Traffic Shaping
    NeTS: Small: Collaborative Research: Distributed Approximate Packet Classification
    海外基金