SeqRepo: A system for managing local collections of biological sequences.

SeqRepo: A system for managing local collections of biological sequences.
复制标题

DOI:
10.1371/journal.pone.0239883
复制
发表时间:
2020
期刊:
影响因子:
3.7
通讯作者:
Prlić A
Prlić A
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hart RK;Prlić A

文献摘要

参考文献

被引文献

相似文献

访问生物序列数据,如基因组、转录本或蛋白质序列,是许多生物信息学分析工作流程的核心。国家生物技术信息中心(NCBI)、Ensembl和其他序列数据库维护者提供了通过网络连接访问序列的方法。对于许多用户来说,远程管理数据的便利性和通用性是引人注目的,并且网络延迟是不重要的。然而,对于高通量和临床应用,局部序列收集对于性能、稳定性、隐私性和再现性至关重要。在这里,我们描述了SeqRepo,一种用于构建本地,高性能,非冗余生物序列集合的新系统。SeqRepo使客户端能够使用主数据库标识符和多个数据库标识符来识别序列和序列别名。SeqRepo提供了一个原生Python接口和一个REST接口,可以在本地运行,并支持从其他编程语言访问。SeqRepo还提供了一个基于GA 4GH refget协议的替代REST接口。SeqRepo提供对序列切片的快速随机访问。我们提供的结果表明,本地SeqRepo序列集合产生显着的性能优势高达1300倍以上的远程序列集合。在我们的变体验证和标准化管道用例中,SeqRepo相对于使用远程序列提高了50倍的吞吐量。SeqRepo可以与任何物种或序列类型一起使用。人类序列收集的定期快照可用。使用计算摘要作为序列标识符通常是方便的或必要的。例如,基于图的标识符可以用于指代专有参考基因组或图基因组的片段,对于这些基因组或片段,常规标识符将不可用。在这里,我们还介绍了一种使用SHA-512哈希算法和Base64编码来生成URL安全标识符的约定。这个约定sha 512 t24 u将快速摘要机制与可用于任何对象的空间高效表示相结合。我们的报告包括对sha 512 t24 u的时间和碰撞概率的分析。SeqRepo使客户端能够使用sha 512 t24 u作为标识符,从而无缝集成公共和私有序列集。SeqRepo在Apache许可证2.0下发布,并可在github和PyPi上使用。Docker镜像和数据库快照也可用。参见https://github.com/biocommons/biocommons.seqrepo。
Access to biological sequence data, such as genome, transcript, or protein sequence, is at the core of many bioinformatics analysis workflows. The National Center for Biotechnology Information (NCBI), Ensembl, and other sequence database maintainers provide methods to access sequences through network connections. For many users, the convenience and currency of remotely managed data are compelling, and the network latency is non-consequential. However, for high-throughput and clinical applications, local sequence collections are essential for performance, stability, privacy, and reproducibility. Here we describe SeqRepo, a novel system for building a local, high-performance, non-redundant collection of biological sequences. SeqRepo enables clients to use primary database identifiers and several digests to identify sequences and sequence alises. SeqRepo provides a native Python interface and a REST interface, which can run locally and enables access from other programming languages. SeqRepo also provides an alternative REST interface based on the GA4GH refget protocol. SeqRepo provides fast random access to sequence slices. We provide results that demonstrate that a local SeqRepo sequence collection yields significant performance benefits of up to 1300-fold over remote sequence collections. In our use case for a variant validation and normalization pipeline, SeqRepo improved throughput 50-fold relative to use with remote sequences. SeqRepo may be used with any species or sequence type. Regular snapshots of Human sequence collections are available. It is often convenient or necessary to use a computed digest as a sequence identifier. For example, a digest-based identifier may be used to refer to proprietary reference genomes or segments of a graph genome, for which conventional identifiers will not be available. Here we also introduce a convention for the application of the SHA-512 hashing algorithm with Base64 encoding to generate URL-safe identifiers. This convention, sha512t24u, combines a fast digest mechanism with a space-efficient representation that can be used for any object. Our report includes an analysis of timing and collision probabilities for sha512t24u. SeqRepo enables clients to use sha512t24u as identifiers, thereby seamlessly integrating public and private sequence sets. SeqRepo is released under the Apache License 2.0 and is available on github and PyPi. Docker images and database snapshots are also available. See https://github.com/biocommons/biocommons.seqrepo.
DOI: 10.1093/bioinformatics/btv112
发表时间: 2015-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Tan, Adrian;Abecasis, Goncalo R.;Kang, Hyun Min
通讯作者: Kang, Hyun Min
DOI: 10.1002/humu.22981
发表时间: 2016-06-01
期刊: HUMAN MUTATION
影响因子: 3.9
作者:
den Dunnen, Johan T.;Dalgleish, Raymond;Taschner, Peter E. M.
通讯作者: Taschner, Peter E. M.
DOI: 10.1093/database/bax020
发表时间: 2017-01-01
期刊: Database : the journal of biological databases and curation
影响因子: --
作者:
Ruffier M;Kähäri A;Komorowska M;Keenan S;Laird M;Longden I;Proctor G;Searle S;Staines D;Taylor K;Vullo A;Yates A;Zerbino D;Flicek P
通讯作者: Flicek P
DOI: 10.1002/humu.23615
发表时间: 2018-12
期刊: Human mutation
影响因子: 3.9
作者:
Wang M;Callenberg KM;Dalgleish R;Fedtsov A;Fox NK;Freeman PJ;Jacobs KB;Kaleta P;McMurry AJ;Prlić A;Rajaraman V;Hart RK
通讯作者: Hart RK
DOI: 10.1093/bioinformatics/btq671
发表时间: 2011-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, Heng
通讯作者: Li, Heng