课题基金 / 基金详情

Compute and Storage Cluster

Compute and Storage Cluster
计算和存储集群
批准号:
469073465
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Major Research Instrumentation
财政年份:
2021
资助国家:
德国
项目状态:
未结题
起止时间:
2020-12-31 至 --
关键词:

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
In order to store and efficiently process data from modern high-throughput methods, a memory and a CPU compute cluster are required, among other things. The cluster applied for will mainly store and process metagenomic data (the entire genome of all bacteria in a given sample) and single cell research data (single-cell RNA seq and spatial single cell transcriptomics). These two data types are currently among the largest volume types in the molecular data environment. The original compute cluster acquired in 2013 is designed to process molecular data obtained from microarray technology. This technology is now no longer used in the participating groups and has been completely replaced by sequencing. However, sequencing data requires significantly more memory and computational effort. Together with its partners, including the Helmholtz Institute for Pharmaceutical Research Saarland (HIPS) and several departments of Saarland University Hospital (UKS), the bioinformatics department generates about 2,000 metagenomes per year with a depth of 15 gigabases per sample. The human genetics and bioinformatics departments of Saarland University, together with their partners, generate and process an additional approximately 2 million RNA single-cell profiles annually using so-called Drop-seq methods. In pilot projects, ATAC-seq is currently performed in addition to RNA-seq and RNA profiles are collected at subcellular resolution. One single-cell experiment with 50,000 cells - sequenced in 2 days - requires about 3 TB of storage and 20 days of pure primary data analysis time. During the analysis of the data, intermediate results are sometimes generated that are more comprehensive than the primary data itself. The requested large-scale system should be able to store at least 100 experiments in parallel and redundantly, reducing the processing time from 20 days to about 3 days. To achieve this, the system requires at least 1,700 TB of gross storage capacity (for example, 6 x 16 x 18 GB HDDs) and at least 512 computational cores (for example, 16 x 32-core processors) with a clock frequency between 2.5 and 3 GHz. Since the analyses performed are generally memory-intensive, 8 TB of RAM should be available. A decisive factor is to avoid so-called swapping, i.e. the frequent copying of data between the RAM and the hard disk. Therefore, a total of at least 16 TB of buffer memory and 100 TB of fast data storage (solid state disks; SSDs) for all processors together are required. Furthermore, a 100 Gb network - consisting of network cards and a corresponding Gb switch - is essential so that copying the data does not become a bottleneck. In addition, a so-called metadata server is needed, which can optimally distribute jobs and processes to the individual components. The metadata server should be equipped with four 32-core processors and 2 TB RAM.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
面向 In-Storage 智能计算的高性能 SSD 控制器研究
  • 批准号:
    ZCLJHSQY26F0401
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    何越
  • 依托单位:
面向in-storage智能计算的固态硬盘缓存管理优化
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2022
  • 负责人:
    廖剑伟
  • 依托单位: