NEARDATA - Extreme Near-Data Processing Platform
NEARDATA - Extreme Near-Data Processing Platform
批准号:
10048448
负责人:
金额:
$39.53万
依托单位国家:
英国
项目类别:
EU-Funded
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
其主要目标是设计一个极接近数据的平台,以支持分布式和联合数据的消费、挖掘和处理,而不需要掌握跨不同数据位置和池的数据访问的物流。我们超越了从存储系统接收的传统被动或批量数据,转向了云和边缘中的新一代近数据处理平台。在我们的平台中,极端数据包括元数据和可信数据连接器,支持高级数据管理操作,如从异类数据源进行数据发现、挖掘和过滤。三个核心目标是:O-1为极端数据类型提供高性能的近数据处理:第一个目标是创建一个新的中间数据服务(XtremeDataHub),提供无服务器数据连接器,优化数据管理操作(分区、过滤、转换、聚合)和交互式查询(搜索、发现、匹配、多对象查询),以高效地向分析平台呈现数据。我们的数据连接器促进了ELA-tic数据驱动的流程-然后计算模式,这显著减少了数据互连上的数据通信,最终导致更高的总体数据吞吐量。O-2支持实时视频流,但也支持必须以非常快的速度接收和处理到对象存储的事件流:第二个目标是无缝结合流和批处理数据处理进行分析。为此,我们将开发部署为流运算符的流数据连接器,在低延迟事件和视频流上提供非常快速的有状态计算。O-3第三个目标是创建一个数据代理服务,以实现可信的数据共享和跨计算连续统的数据管道的机密编排。借助可信执行环境(TEE)和联合学习架构,我们将提供安全的数据协调、传输、处理和访问
英文摘要
The main goal is to design an Extreme near-data platform to enable consumption, mining and processing of dis- tributed and federated data without needing to master the logistics of data access across heterogeneous data locations and pools. We go beyond traditional passive or bulk data ingested from storage systems towards next generation near-data processing platforms both in the Cloud and in the Edge. In our platform, Extreme Data in- cludes both metadata and trustworthy data connectors enabling advanced data management operations like data discovery, mining, and filtering from heterogeneous data sources. The three core objectives are: O-1 Provide high-performance near-data processing for Extreme Data Types: The first objective is to create a novel intermediary data service (XtremeDataHub) providing serverless data connectors that optimize data management operations (partitioning, filtering, transformation, aggregation) and interactive queries (search, discovery, matching, multi-object queries) to efficiently present data to analytics platforms. Our data connectors facilitate a elas- tic data-driven process-then-compute paradigm which significantly reduces data communication on the data interconnect, ultimately resulting in higher overall data throughput. O-2 Support real-time video streams but also event streams that must be ingested and processed very fast to Object Storage: The second objective is to seamlessly combine streaming and batch data processing for analytics. To this end, we will develop stream data connectors deployed as stream operators offering very fast stateful computations over low-latency event and video streams. O-3 The third objective is to create a Data Broker service enabling trustworthy data sharing and confidential orchestration of data pipelines across the Compute Continuum. We will provide secure data orchestration, transfer, processing and access thanks to Trusted Execution Environments (TEEs) and federated learning architectures
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金