课题基金 / 基金详情

Compute and Storage Cluster

Compute and Storage Cluster
计算和存储集群
批准号:
469073465
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Major Research Instrumentation
财政年份:
2021
资助国家:
德国
项目状态:
未结题
起止时间:
2020-12-31 至 --
关键词:

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
为了存储和有效地处理来自现代高吞吐量方法的数据,除其他外,还需要存储器和CPU计算集群。申请的集群将主要存储和处理宏基因组数据(给定样本中所有细菌的整个基因组)和单细胞研究数据(单细胞RNA seq和空间单细胞转录组学)。这两种数据类型目前是分子数据环境中最大的卷类型。2013年获得的原始计算集群旨在处理从微阵列技术获得的分子数据。这项技术现在已不再用于参与组,并已完全被测序取代。然而,测序数据需要显著更多的存储器和计算工作。与其合作伙伴,包括亥姆霍兹药物研究所萨尔兰(HIPS)和萨尔兰大学医院(UKS)的几个部门一起,生物信息学部门每年产生约2,000个宏基因组,每个样本的深度为15千兆碱基。萨尔兰大学的人类遗传学和生物信息学系与他们的合作伙伴一起,每年使用所谓的Drop-seq方法生成和处理额外的约200万个RNA单细胞图谱。在试点项目中,除了RNA-seq之外,目前还进行ATAC-seq,并以亚细胞分辨率收集RNA图谱。一个50,000个细胞的单细胞实验-在2天内测序-需要大约3 TB的存储和20天的纯原始数据分析时间。在数据分析过程中,有时会产生比原始数据本身更全面的中间结果。所要求的大规模系统应该能够并行和冗余地存储至少100个实验,将处理时间从20天减少到大约3天。为了实现这一目标,系统需要至少1,700 TB的总存储容量(例如,6 x 16 x 18 GB HDD)和至少512个计算核心(例如,16 x 32核处理器),时钟频率在2.5和3 GHz之间。由于执行的分析通常是内存密集型的,因此应该有8 TB的RAM可用。一个决定性因素是避免所谓的交换,即在RAM和硬盘之间频繁复制数据。因此,所有处理器总共需要至少16 TB的缓冲存储器和100 TB的快速数据存储(固态硬盘; SSD)。此外,由网卡和相应的Gb交换机组成的100 Gb网络至关重要,这样复制数据就不会成为瓶颈。此外,还需要一个所谓的元数据服务器,它可以最佳地将作业和进程分配给各个组件。元数据服务器应配备4个32核处理器和2 TB RAM。
英文摘要
In order to store and efficiently process data from modern high-throughput methods, a memory and a CPU compute cluster are required, among other things. The cluster applied for will mainly store and process metagenomic data (the entire genome of all bacteria in a given sample) and single cell research data (single-cell RNA seq and spatial single cell transcriptomics). These two data types are currently among the largest volume types in the molecular data environment. The original compute cluster acquired in 2013 is designed to process molecular data obtained from microarray technology. This technology is now no longer used in the participating groups and has been completely replaced by sequencing. However, sequencing data requires significantly more memory and computational effort. Together with its partners, including the Helmholtz Institute for Pharmaceutical Research Saarland (HIPS) and several departments of Saarland University Hospital (UKS), the bioinformatics department generates about 2,000 metagenomes per year with a depth of 15 gigabases per sample. The human genetics and bioinformatics departments of Saarland University, together with their partners, generate and process an additional approximately 2 million RNA single-cell profiles annually using so-called Drop-seq methods. In pilot projects, ATAC-seq is currently performed in addition to RNA-seq and RNA profiles are collected at subcellular resolution. One single-cell experiment with 50,000 cells - sequenced in 2 days - requires about 3 TB of storage and 20 days of pure primary data analysis time. During the analysis of the data, intermediate results are sometimes generated that are more comprehensive than the primary data itself. The requested large-scale system should be able to store at least 100 experiments in parallel and redundantly, reducing the processing time from 20 days to about 3 days. To achieve this, the system requires at least 1,700 TB of gross storage capacity (for example, 6 x 16 x 18 GB HDDs) and at least 512 computational cores (for example, 16 x 32-core processors) with a clock frequency between 2.5 and 3 GHz. Since the analyses performed are generally memory-intensive, 8 TB of RAM should be available. A decisive factor is to avoid so-called swapping, i.e. the frequent copying of data between the RAM and the hard disk. Therefore, a total of at least 16 TB of buffer memory and 100 TB of fast data storage (solid state disks; SSDs) for all processors together are required. Furthermore, a 100 Gb network - consisting of network cards and a corresponding Gb switch - is essential so that copying the data does not become a bottleneck. In addition, a so-called metadata server is needed, which can optimally distribute jobs and processes to the individual components. The metadata server should be equipped with four 32-core processors and 2 TB RAM.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
面向 In-Storage 智能计算的高性能 SSD 控制器研究
  • 批准号:
    ZCLJHSQY26F0401
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    何越
  • 依托单位:
面向in-storage智能计算的固态硬盘缓存管理优化
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2022
  • 负责人:
    廖剑伟
  • 依托单位: