Named Data Networking for Genomics Data Management and Integrated Workflows.

Named Data Networking for Genomics Data Management and Integrated Workflows.
复制标题

命名为基因组数据管理和集成工作流的数据网络。

DOI:
10.3389/fdata.2021.582468
复制
发表时间:
2021
影响因子:
3.1
通讯作者:
Shannigrahi S
Shannigrahi S
中科院分区:
其他
文献类型:
--
作者:
Ogle C;Reddick D;McKnight C;Biggs T;Pauly R;Ficklin SP;Feltus FA;Shannigrahi S

文献摘要

参考文献

被引文献

相似文献

先进的成像和DNA测序技术现在使不同的生物界能够例行公事地生成和分析TB级的高分辨率生物数据。在单个研究人员的实验室环境中,社区正在迅速走向千万亿级。作为证据,单一的NCBI SRA中央DNA序列存储库包含超过45 PB的生物数据。鉴于这个和其他基因组库的几何级增长,一艾字节的可挖掘生物学数据迫在眉睫。有效利用这些数据集的挑战是巨大的,因为它们不仅大,而且还存储在不同储存库中的地理分布储存库中,如国家生物技术信息中心(NCBI)、日本DNA数据库(DDBJ)、欧洲生物信息学研究所(EBI)和美国国家航空航天局(NASA)的基因实验室。在这项工作中,我们首先系统地指出了基因组学社区的数据管理挑战。然后,我们介绍了命名数据网络(NDN),这是一种新颖但研究得很好的互联网体系结构,能够在网络层解决这些挑战。NDN使用内容名称(类似于传统的文件名或文件路径)执行所有操作,如将请求转发到数据源、内容发现、访问和检索,并且不再需要用于数据管理的位置层(IP地址)。将NDN用于基因组学工作流简化了数据发现,使用流行数据集的网络内缓存加快了数据检索,并允许社区创建支持诸如创建内容存储库联合、从多个来源检索、远程数据子设置等操作的基础架构。基于命名的操作还简化了工作流与各种云平台的部署和集成。我们在这项工作中的贡献如下:1)我们列举了NDN可以缓解的基因组学社区的网络基础设施挑战,2)我们描述了我们在将NDN应用于当代基因组学工作流(GEMMaker)方面所做的努力,并量化了改进。初步评估显示,将数据插入工作流程的速度提高了六倍。3)作为试点,我们使用NDN命名方案(由社区同意,并在第4节中讨论)发布来自广泛使用的数据存储库的数据,包括NCBI SRA。我们已经在NDN试验台上加载了这些经过预处理的基因组,这些基因组可以通过NDN访问,任何对这些数据集感兴趣的人都可以使用。最后,我们讨论了我们在将NDN与云计算平台(如Pacific Research Platform(PRP))集成方面所做的持续努力。读者应该注意到,这篇文章的目的是将NDN介绍给基因组学社区,并讨论NDN的特性可以使基因组学社区受益。我们没有提出对NDN的广泛业绩评估--我们正在努力扩大和评估我们的试点部署,并将在未来的工作中提出系统的结果。
Advanced imaging and DNA sequencing technologies now enable the diverse biology community to routinely generate and analyze terabytes of high resolution biological data. The community is rapidly heading toward the petascale in single investigator laboratory settings. As evidence, the single NCBI SRA central DNA sequence repository contains over 45 petabytes of biological data. Given the geometric growth of this and other genomics repositories, an exabyte of mineable biological data is imminent. The challenges of effectively utilizing these datasets are enormous as they are not only large in the size but also stored in geographically distributed repositories in various repositories such as National Center for Biotechnology Information (NCBI), DNA Data Bank of Japan (DDBJ), European Bioinformatics Institute (EBI), and NASA’s GeneLab. In this work, we first systematically point out the data-management challenges of the genomics community. We then introduce Named Data Networking (NDN), a novel but well-researched Internet architecture, is capable of solving these challenges at the network layer. NDN performs all operations such as forwarding requests to data sources, content discovery, access, and retrieval using content names (that are similar to traditional filenames or filepaths) and eliminates the need for a location layer (the IP address) for data management. Utilizing NDN for genomics workflows simplifies data discovery, speeds up data retrieval using in-network caching of popular datasets, and allows the community to create infrastructure that supports operations such as creating federation of content repositories, retrieval from multiple sources, remote data subsetting, and others. Named based operations also streamlines deployment and integration of workflows with various cloud platforms. Our contributions in this work are as follows 1) we enumerate the cyberinfrastructure challenges of the genomics community that NDN can alleviate, and 2) we describe our efforts in applying NDN for a contemporary genomics workflow (GEMmaker) and quantify the improvements. The preliminary evaluation shows a sixfold speed up in data insertion into the workflow. 3) As a pilot, we have used an NDN naming scheme (agreed upon by the community and discussed in Section 4) to publish data from broadly used data repositories including the NCBI SRA. We have loaded the NDN testbed with these pre-processed genomes that can be accessed over NDN and used by anyone interested in those datasets. Finally, we discuss our continued effort in integrating NDN with cloud computing platforms, such as the Pacific Research Platform (PRP). The reader should note that the goal of this paper is to introduce NDN to the genomics community and discuss NDN’s properties that can benefit the genomics community. We do not present an extensive performance evaluation of NDN—we are working on extending and evaluating our pilot deployment and will present systematic results in a future work.
DOI: 10.1177/1177932219856359
发表时间: 2019-06-14
影响因子: 5.8
作者:
Mills, Nicholas;Bensman, Ethan M.;Feltus, F. Alex
通讯作者: Feltus, F. Alex
DOI: 10.1038/s41598-017-09094-4
发表时间: 2017-08-17
期刊: Scientific reports
影响因子: 4.6
作者:
Ficklin SP;Dunwoodie LJ;Poehlman WL;Watson C;Roche KE;Feltus FA
通讯作者: Feltus FA
DOI: 10.1093/nar/gkp1137
发表时间: 2010-04
影响因子: 14.9
作者:
Cock PJ;Fields CJ;Goto N;Heuer ML;Rice PM
通讯作者: Rice PM
DOI: 10.1109/jproc.2009.2021005
发表时间: 2009-08-01
影响因子: 20.6
作者:
Dewdney, Peter E.;Hall, Peter J.;Lazio, T. Joseph L. W.
通讯作者: Lazio, T. Joseph L. W.
DOI: 10.1186/1471-2105-12-361
发表时间: 2011-09-09
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Chiang, Gen-Tao;Clapham, Peter;Coates, Guy
通讯作者: Coates, Guy