课题基金 / 基金详情

Context-based Pattern Recognition for Automating Big Data Management in Clouds

Context-based Pattern Recognition for Automating Big Data Management in Clouds
基于上下文的模式识别,用于自动化云中的大数据管理
批准号:
RGPIN-2017-06925
负责人:
Amza, Cristiana
金额:
$1.89万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Amza, Cristiana的其他基金

相似基金

相关文献

中文摘要
翻译
我们建议在云环境产生的大量操作数据中构建实时、动态模式识别的数据模型。我们的模式识别技术将应用于云环境的一系列自动化管理任务,例如异常诊断和资源供应。我们还将研究我们的模式识别技术对其他类型大数据的可扩展性和适用性,例如托管在云环境上的生物医学和生物识别数据。******目标***1。数据收集***我们将捕获各种类型的云操作数据,并在收集点为所有感兴趣的应用程序范围(例如,工作负载阶段、参数化请求或方法调用)标记它们所处的高级应用程序范围。我们计划捕获和调查的操作数据类型包括日常日志记录和监控信息,例如吞吐量、延迟、可靠性或其他应用程序QoS指标,以及日志语句和CPU、网络带宽、磁盘和内存消耗的资源消耗时间序列数据。******2。模式学习和识别***在收集数据时,根据上下文部署一种或多种统计学习技术来学习数学数据模型。对于统计上偏离学习数据模型的上下文,我们将触发自动系统适应,例如,在资源分配方面和/或提供上下文丰富的问题摘要供人类检查。******方法论***我们的关键思想是利用动态模式学习和识别框架中的应用程序上下文,通过显式地增加常规系统日志记录和使用这些数据发生的应用程序上下文监视数据收集。我们期望几种互补的数据建模和分析方法将联合或独立地应用于各种上下文,例如,用于相同的应用范围,用于相同粒度的不同范围,和/或用于不同粒度的嵌套范围。不同的方法可能更适合不同的数据和不同的上下文。例如,基于聚类的分类方法可以很好地处理日志模板数据,但是基于资源消耗的时间序列数据的模式匹配的神经网络方法可以提供更好的模式提取、学习和匹配精度。这两种方法都比单独使用任何一种方法都能提供更高的异常检测精度。为了检测应用程序上下文,我们目前对应用程序源代码使用静态分析。然而,随着项目的进展,我们计划包括更多基于二进制代码检查的通用技术。在我们对原型的实验评估中,我们将在我们的云软件堆栈中使用开源系统,并实验聚类、动态时间扭曲、神经网络和其他统计学习方法。
英文摘要
We propose to build data models for real-time, on-the-fly pattern recognition in the large volumes of operational data produced by Cloud environments. Our pattern recognition techniques will be applied to a range of automated management tasks for Cloud environments, such as, anomaly diagnosis, and resource provisioning. We will also investigate the extensibility and applicability of our pattern recognition techniques to other types of big data, such as, biomedical and biometric data hosted on Cloud environments.******Objectives***1. Data Collection***We will capture various types of Cloud operational data and tag them with the high level application scope in which they occur, at the collection points, for all the application scopes of interest e.g., workload phase, parametrized request, or method invocation. The types of operational data we plan to capture and investigate include routine logging and monitoring information, such as, throughput, latency, dependability, or other application QoS metrics, as well as log statements and resource consumption time series data for CPU, network bandwidth, disk and memory consumption.******2. Pattern Learning and Recognition***As data is collected, one or more statistical learning techniques are deployed in order to learn mathematical data models, per context. For contexts that statistically diverge from the learned data model, we will trigger an automated system adaptation e.g., in terms of resource allocation and/or provide context-rich problem digests for human inspection.******Methodology***Our key idea is to leverage application contexts within our on-the-fly pattern learning and recognition framework, by explicitly augmenting routine system logging and monitoring data collection with the application contexts where these data occur. We expect that several, complementary data modelling and analysis methods will be applied jointly, or independently, for the various contexts, e.g., for the same application scope, for different scopes of the same granularity, and/or for nested scopes of different granularities. Different methods may work better for different data and different contexts. For example, a classification method based on clustering may work well for log template data, but a Neural Network method for pattern matching in time series data for resource consumption could provide better pattern extraction, learning and matching accuracy. Both methods could provide higher accuracy for anomaly detection than either method alone. In order to detect application contexts we currently use static analysis on application source code. However, we plan on including more general techniques based on binary code inspection as our project progresses. In our experimental evaluation of our prototype, we will use open-source systems in our Cloud software stack and experiment with clustering, dynamic time warp, Neural Networks and other statistical learning methods.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Context-based Pattern Recognition for Automating Big Data Management in Clouds
  • 批准号:
    RGPIN-2017-06925
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.79万
  • 财政年份:
    2021
  • 负责人:
    Amza, Cristiana
  • 依托单位:
Context-based Pattern Recognition for Automating Big Data Management in Clouds
  • 批准号:
    RGPIN-2017-06925
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2020
  • 负责人:
    Amza, Cristiana
  • 依托单位:
Context-based Pattern Recognition for Automating Big Data Management in Clouds
  • 批准号:
    RGPIN-2017-06925
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2019
  • 负责人:
    Amza, Cristiana
  • 依托单位:
Context-based Pattern Recognition for Automating Big Data Management in Clouds
  • 批准号:
    RGPIN-2017-06925
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2017
  • 负责人:
    Amza, Cristiana
  • 依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Incentive and governance schenism study of corporate green washing behavior in China: Based on an integiated view of econfiguration of environmental authority and decoupling logic
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    YU BYUNGJUN
  • 依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
  • 批准号:
    W2433169
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    HAOFEI ZHANG
  • 依托单位:
含Re、Ru先进镍基单晶高温合金中TCP相成核—生长机理的原位动态研究
  • 批准号:
    52301178
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30.00万元
  • 批准年份:
    2023
  • 负责人:
    夏万顺
  • 依托单位: