pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive.

pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive.
复制标题

DOI:
10.12688/f1000research.18676.1
复制
发表时间:
2019-01-01
期刊:
影响因子:
--
通讯作者:
Choudhary, Saket
Choudhary, Saket
中科院分区:
其他
文献类型:
--
作者:
Choudhary, Saket

文献摘要

被引文献

相似文献

NCBI 序列读取存档 (SRA) 是下一代测序数据集的主要存档。 SRA 向研究界提供元数据和原始测序数据,以鼓励可重复性,并为根据公开数据测试新假设提供途径。然而,以编程方式访问这些数据的方法是有限的。我们介绍 Python 包 pysradb,它提供了一系列命令行方法来查询和下载 SRA 中的元数据和数据,利用 SRAdb 项目提供的精选元数据数据库。我们在多个用例中演示了 pysradb 用于搜索和下载 SRA 数据集的实用程序。它可以在 https://github.com/saketkc/pysradb 上免费获得。
The NCBI Sequence Read Archive (SRA) is the primary archive of next-generation sequencing datasets. SRA makes metadata and raw sequencing data available to the research community to encourage reproducibility and to provide avenues for testing novel hypotheses on publicly available data. However, methods to programmatically access this data are limited. We introduce the Python package, pysradb, which provides a collection of command line methods to query and download metadata and data from SRA, utilizing the curated metadata database available through the SRAdb project. We demonstrate the utility of pysradb on multiple use cases for searching and downloading SRA datasets. It is available freely at https://github.com/saketkc/pysradb.