课题基金 / 基金详情

CSR: Small: Empower Data-Intensive Computing: the integrated data management approach

CSR: Small: Empower Data-Intensive Computing: the integrated data management approach
CSR:小:赋能数据密集型计算:集成数据管理方法
批准号:
1526887
负责人:
Xian-He Sun
金额:
$40.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2019-08-31

项目摘要

项目成果

Xian-He Sun的其他基金

相似基金

相关文献

中文摘要
翻译
从计算机系统的角度来看,有两种类型的数字数据:观测数据,由传感器、监视器、相机、文本等电子设备收集的数据;以及仿真数据,通过计算产生的数据。前者代表新兴的互联网数据驱动应用,如社交媒体和数据分析;后者代表传统计算驱动的应用,如气候建模和计算流体动力学。一般来说,后者需要很强的一致性才能保证正确性,而前者则不需要。一致性的不同导致了两种文件系统:数据密集型分布式文件系统,以基于MapReduce的Hadoop分布式文件系统(HDFS)为代表;计算密集型文件系统,以高性能并行文件系统(PFS)为代表,如IBM通用并行文件系统(GPFS)。这两种文件系统以不同的设计理念设计,用于不同的应用程序,并且不相互通信。理解大量收集的数据依赖于强大的计算,而大规模计算需要管理大量数据。因此,大数据应用需要一个一体化的解决方案。本研究开发的集成数据访问系统(IDAS)旨在弥合数据管理的鸿沟,遵循分布式系统设计中的CAP理论,IDAS方法不是作为一个新的独立系统设计,而是作为一个软件层,提供一个集成接口,在不改变用户应用的情况下,进行跨平台的数据访问,从HDFS到PFS,或从PFS到HDFS,读或写,有效和可互换。IDAS的开发计划包括三个部分:1)建立通信通道,使HDFS和PFS之间可以访问数据;2)设计扩展的语义接口,使不同的计算系统可以访问不同的文件系统;3)开发优化技术,优化HDFS、PFS和IDAS下的I/O操作。大数据需要数据驱动型互联网计算界和计算驱动型科学计算界的共同努力。国际开发协会为跨平台、跨社区的数据存储、访问和共享服务提供了一个可持续、具有成本效益的基础设施。这项研究将创建先进的解决方案和技术,这些解决方案和技术将对提高大规模数据访问和管理的效率产生直接影响。由于大数据是科学、工程和工业的国家战略基础设施,拟议的调查将推进广泛的领域。这项研究的成功将努力在一个及时、重要、极具挑战性和高影响力的问题--综合数据访问系统--方面取得重大进展。
英文摘要
From the computer system point of view there are two types of digital data: observational data, the data collected by electrical devices such as sensor, monitor, camera, text, etc.; and simulation data, data generated by computing. The former represents newly emerged internet data-driven applications, such as social media and data analytic; and the latter represents the conventional computing-driven applications, such as climate modeling and computational fluid dynamics. In general, the latter requires strong consistency for correctness and the former does not. The difference in consistency leads to two kinds of file systems: data-intensive distributed file system, represented by the MapReduce-based Hadoop distributed file systems (HDFS); and computing-intensive file systems, represented by the high performance parallel file systems (PFS), such as the IBM general parallel file system (GPFS). These two kinds of file systems are designed with different philosophies, for different applications, and do not talk to each other. Understanding huge amounts of collected data depends on powerful computation, whereas large-scale computation requires the management of large data. Therefore, big data applications demand an integrated solution. The integrated data access system (IDAS) developed under this research is designed to bridge the data management gap.In agreement with the CAP theory in the distributed system design, the IDAS approach is not designed as a new standalone system but as a software layer which provides an integrated interface to conduct cross-platform data access, from HDFS to PFS, or from PFS to HDFS, read or write, effectively and interchangeably without changing the users' applications. The development plan for IDAS has three components: 1) establish the communication channels so that data can be accessed between HDFS and PFS; 2) design an extended semantic interface so that different file systems can be accessed under different computing systems; 3) develop optimization techniques to optimize I/O operation under HDFS, PFS, and under IDAS. Big data requires a joint effort of the data-driven internet computing community and the compute-driven scientific computing community. IDAS provides a sustainable, cost-effective infrastructure for cross-platform, cross-community services of data storage, access, and sharing. This research will create advanced solutions and technologies that will have direct impact on improving the efficiency of data access and management at scale. Since big data is a national strategic infrastructure for science, engineering, and industry, the proposed investigations will advance a broad range of fields. The success of this research will strive to make significant progress of a timely, important, highly challenging, and high-impact problem, namely integrated data access system.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
OAC Core: LABIOS: Storage Acceleration via Data Labeling and Asynchronous I/O
  • 批准号:
    2313154
  • 项目类别:
    Standard Grant
  • 资助金额:
    $60.0万
  • 财政年份:
    2023
  • 负责人:
    Xian-He Sun
  • 依托单位:
Collaborative Research: CSR: Medium: Towards A Unified Memory-centric Computing System with Cross-layer Support
  • 批准号:
    2310422
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $75.0万
  • 财政年份:
    2023
  • 负责人:
    Xian-He Sun
  • 依托单位:
CNS Core: Small: Practical Memory Access Pattern Obfuscation with Algorithm, Application and Architecture Co-designs
  • 批准号:
    2152497
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.8万
  • 财政年份:
    2022
  • 负责人:
    Xian-He Sun
  • 依托单位:
Frameworks: Collaborative Research: ChronoLog: A High-Performance Storage Infrastructure for Activity and Log Workloads
  • 批准号:
    2104013
  • 项目类别:
    Standard Grant
  • 资助金额:
    $267.65万
  • 财政年份:
    2021
  • 负责人:
    Xian-He Sun
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: