CNSA: a data repository for archiving omics data

CNSA: a data repository for archiving omics data
复制标题

DOI:
10.1093/database/baaa055
复制
发表时间:
2020-07-23
影响因子:
5.8
通讯作者:
Xu, Xun
Xu, Xun
中科院分区:
生物学4区
文献类型:
--
作者:
Guo, Xueqin;Chen, Fengzhen;Xu, Xun

文献摘要

被引文献

相似文献

随着高通量测序技术在生命和健康科学中的应用和发展,海量的多组学数据带来了高效管理和利用的问题。数据库开发和生物修复是这些大数据再利用的前提。在此,我们依托中国国家基因库,提出了一种用于组学数据归档的CNGB序列档案库,包括原始测序数据及其进一步分析的结果,这些数据被组织成目前的项目、样本、实验、运行、组装和变异六个对象。此外,国家航天局还建立了一些项目的活样本、样本信息和分析数据的关联模型。活体样品和分析数据都与样品信息直接相关。从其中一个可以获得另外两个的信息或数据,从而可以追溯从活样本到样本信息再到分析数据的整个生命周期的所有数据。遵循生命科学中常用的数据标准,CNSA致力于建立一个全面的、经过精心策划的数据存储库,用于存储、管理和共享组学数据。我们将继续完善数据标准,为全球科学界免费提供开放数据资源,以支持学术研究和生物产业。
With the application and development of high-throughput sequencing technology in life and health sciences, massive multi-omics data brings the problem of efficient management and utilization. Database development and biocuration are the prerequisites for the reuse of these big data. Here, relying on China National GeneBank (CNGB), we present CNGB Sequence Archive (CNSA) for archiving omics data, including raw sequencing data and its further analyzed results which are organized into six objects, namely Project, Sample, Experiment, Run, Assembly and Variation at present. Moreover, CNSA has created a correlation model of living samples, sample information and analytical data on some projects. Both living samples and analytical data are directly correlated with the sample information. From either one, information or data of the other two can be obtained, so that all data can be traced throughout the life cycle from the living sample to the sample information to the analytical data. Complying with the data standards commonly used in the life sciences, CNSA is committed to building a comprehensive and curated data repository for storing, managing and sharing of omics data. We will continue to improve the data standards and provide free access to open-data resources for worldwide scientific communities to support academic research and the bio-industry.