DIANA: Data Interface All-iN-A-place for Big Data

DIANA: Data Interface All-iN-A-place for Big Data
复制标题

DIANA:大数据的数据接口 All-iN-A-place

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Frank Z. Wang
Frank Z. Wang
中科院分区:
--
文献类型:
--
作者:
Frank Z. Wang

文献摘要

被引文献

相似文献

大数据中的“多样性”意味着我们拥有广泛的数据类型和来源:例如,文件系统和数据库系统作为两种流行的数据访问接口共存了几十年。本文的工作是通过提出一个数据接口All-iN-A-place(DIANA)来统一这两个接口。第一个挑战在于区分结构化和非结构化数据,并将它们转移到不同的底层平台。结果表明,在索引中的5000的加速已经实现了在提取属性的速度减慢100的代价。构建了一个基于DIANA的云存储系统,用于多功能、远距离、大容量的大数据访问操作,以解决大数据中的“Volume”和“Velocity”问题。它在套接字层封装了一个动态的多流/多路径引擎,符合可移植操作系统接口(POSIX)。
“Variety” in Big Data means we have a wide range of data types and sources: e.g. file systems and database systems co-exist for decades as two popular data-accessing interfaces. This work is to unify these two interfaces by presenting a Data Interface All-iN-A-place (DIANA). The first challenge lies in distinguishing structured and un-structured data and diverting them to different underlying platforms. It is demonstrated that a speedup of 5000 in indexing has been achieved at the expense of a slowdown of 100 in extracting attributes. A DIANA-based cloud storage system is constructed for versatile, long distance and large volume big data accessing operations to address “Volume” and “Velocity” in Big Data. It encapsulates a dynamic multi-stream/multi-path engine at the socket level, which conforms to Portable Operating System Interface (POSIX).