课题基金 / 基金详情

Online Learning in Big-Data Stream Mining

Online Learning in Big-Data Stream Mining
大数据流挖掘在线学习
批准号:
1407712
负责人:
Ali Sayed
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-08-31

项目摘要

项目成果

Ali Sayed的其他基金

相似基金

相关文献

中文摘要
翻译
当今世界日益成为信息驱动的世界。大量的数据来自不同的来源,以不同的格式产生,包括传感器读数、生理测量、文档、电子邮件、交易、推文以及音频或视频文件。许多企业和政府机构也在拥抱自动化,并依靠各种传感器和基础设施来持续收集、存储和分析数据。为了更好地管理物理系统、得出明智的决策、调整生产流程和优化物流选择,赋予评估系统实时处理来自传感器的流信息的能力变得至关重要。数据流挖掘指的是一大类技术,可用于连续接收来自多个来源的数据流的感知和响应系统,并采用旨在检测和预测可操作信息的分析。这些技术在许多领域都很有用,包括医疗和健康信息学,交通、安全和能源的智能连接网络系统,以及社会、多媒体和商业智能。本提案的目的是开发实时流挖掘的方法和技术,目的是从大数据流中提取信息。该框架通过将多个学习者联网,在不同层次实时地提出和回答查询,从而实现了这一目标。该框架的一个核心方面是它支持分布式数据源和数据的分布式处理。除了上述的应用之外,这项研究预计将对用户界面、人机交互、机器对机器通信和服务产生影响。本研究的重点是开发一个框架,使用部署在分布式计算基础设施上的自适应学习器/分类器网络,从大容量数据流中提取分布式知识。所提出的范式不同于现有的主要由查询驱动的挖掘和搜索解决方案。相反,所提出的框架是数据和概念驱动的,可以通过允许网络学习者根据延迟、资源和数据特征进行持续学习和动态适应,并允许各种学习者根据自己的能力和知识主动推理和塑造与其他学习者的互动,从而催化网络流挖掘应用程序的设计和实现的转变。在这方面,本提案中开发的流挖掘方法解决了几个独特的技术挑战:(a)需要开发分散的流挖掘方法,其中学习者根据与其邻居的本地交互做出决策。此步骤包括正式定义本地目标和指标,以及相关的节点间消息交换,从而能够将应用程序分解为一组自主操作的节点,同时确保全局性能;(b)需要开发算法,以应付异步事件,包括节点的不同数据速率、链路故障和动态拓扑配置;(c)由于庞大的数据量和有限的系统资源(包括中央处理器、内存和输入/输出带宽),需要发展分布式解决方案,以有效应付系统过载。每个学习器通常会产生大量的计算成本,并且解决方案需要对单个学习器处理数据的速度敏感;(d)需要开发自适应流挖掘系统来跟踪概念漂移,特别是由于许多因素,包括共享处理节点的拥塞和处理节点之间的通信延迟,数据特征随着时间的推移而演变。
英文摘要
Online Learning in Big Data Stream MiningThe world is increasingly information-driven. Vast amounts of data are being produced by diverse sources and in diverse formats including sensor readings, physiological measurements, documents, emails, transactions, tweets, and audio or video files. Many businesses and government institutions are also embracing automation and relying on a variety of sensors and infrastructure to collect, store, and analyze data on a continuous basis. It is becoming critical to endow assessment systems with the ability to process streaming information from sensors in real-time in order to better manage physical systems, derive informed decisions, tweak production processes, and optimize logistics choices. Data stream mining refers to the broad class of techniques that can be used in sense and respond systems that continuously receive data streams from multiple sources and employ analytics aimed at detecting and predicting actionable information. Such techniques are useful in many domains including medical and health informatics, intelligent connected network systems for transportation, security, and energy, as well as social, multimedia, and business intelligence. The aim of this proposal is to develop methods and techniques for real-time stream mining with the aim of extracting information from large data streams. The framework accomplishes this objective by networking multiple learners to pose and answer queries at different levels and in real time. A central aspect of the framework is that it accommodates distributed data sources and distributed processing of the data. Besides the aforementioned applications, the proposed research is expected to have an impact on user interfaces, human computer interactions, and machine-to-machine communication and services.This research focuses on developing a framework for distributed knowledge extraction from high-volume data streams using a network of adaptive learners/classifiers that is deployed over a distributed computing infrastructure. The proposed paradigm differs from existing mining and search solutions, which are mainly query driven. Instead, the proposed framework is data and concept driven and can catalyze a shift in the design and implementation of networked stream mining applications by allowing continuous learning and dynamic adaptation of networked learners in response to latency, resource and data characteristics, and by allowing various learners to proactively reason and shape their interactions with other learners based on their capabilities and knowledge. In this regard, the approach to stream mining developed in this proposal addresses several unique technical challenges: (a) the need to develop decentralized approaches for stream mining where learners make decisions based on local interactions with their neighbors. This step involves formally defining local objectives and metrics and associated inter-node message exchanges that enable the decomposition of the application into a set of autonomously operating nodes, while ensuring global performance; (b) the need to develop algorithms that are able to cope with asynchronous events including different data rates at the nodes, link failures, and dynamic topology configurations; (c) the need to develop distributed solutions that can cope effectively with system overload, due to large data volumes and limited system resources (including CPU, memory, and I/O bandwidth). There is usually a large computational cost incurred by each learner and solutions need to be sensitive to the rates at which individual learners can handle data; and (d) the need to develop adaptive stream-mining systems to track concept drifts especially since data characteristics evolve over time due to many factors including congestion at shared processing nodes and communication delays between processing nodes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CIF: Small: Inference over Asymmetric Network and Data Structures
CIF: Large: Collaborative Research: Cooperation and Learning Over Cognitive Networks
CIF: SMALL: Explorations and Insights into Adaptive Networks, Animal Flocking Behavior, and Swarm Intelligence
NSF Workshop on Distributed Processing over Cognitive Networks
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: