Low-Bandwidth and Non-Compute Intensive Remote Identification of Microbes from Raw Sequencing Reads

Low-Bandwidth and Non-Compute Intensive Remote Identification of Microbes from Raw Sequencing Reads
复制标题

DOI:
10.1371/journal.pone.0083784
复制
发表时间:
2013-12-31
期刊:
影响因子:
3.7
通讯作者:
Lund, Ole
Lund, Ole
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Gautier, Laurent;Lund, Ole

文献摘要

被引文献

相似文献

廉价的DNA测序可能很快就会成为常规,不仅适用于人类基因组,而且适用于几乎任何需要从DNA中识别生物体的事情:追踪传染病病原体,控制食品,生物反应器或环境样本。我们提出了一种新的通用方法来分析测序数据,其中不需要指定参考基因组。使用分布式架构,我们能够查询远程服务器,以获得有关引用可能是什么的提示,传输相对少量的数据。我们的系统包括一个服务器与已知的参考DNA索引,和一个客户端与原始测序读取。客户端发送一个未识别读取的样本,并返回一个匹配引用的列表。可以检索参考的序列,并将其用于读取的穷举计算,例如比对。为了证明这种方法,我们已经实现了一个网络服务器,索引数以万计的公开可用的基因组和基因组区域从各种生物体和查询测序读取匹配命中返回列表。我们还实现了两个客户端:一个在Web浏览器中运行,另一个作为Python脚本运行。两者都能够处理来自便携式设备(在平板电脑上运行的基于浏览器的)的大量测序读数,在几秒钟内执行其任务,并消耗与移动的宽带网络兼容的带宽量。这种客户端-服务器方法可以在未来发展,允许完全自动化处理测序数据和台式测序仪测序运行的常规即时质量检查。可通过http://tapir.cbs.dtu.dk访问网站。python命令行客户端、服务器和补充数据的源代码可以在http://bit.ly/1aURxkc上找到。
Cheap DNA sequencing may soon become routine not only for human genomes but also for practically anything requiring the identification of living organisms from their DNA: tracking of infectious agents, control of food products, bioreactors, or environmental samples. We propose a novel general approach to the analysis of sequencing data where a reference genome does not have to be specified. Using a distributed architecture we are able to query a remote server for hints about what the reference might be, transferring a relatively small amount of data. Our system consists of a server with known reference DNA indexed, and a client with raw sequencing reads. The client sends a sample of unidentified reads, and in return receives a list of matching references. Sequences for the references can be retrieved and used for exhaustive computation on the reads, such as alignment. To demonstrate this approach we have implemented a web server, indexing tens of thousands of publicly available genomes and genomic regions from various organisms and returning lists of matching hits from query sequencing reads. We have also implemented two clients: one running in a web browser, and one as a python script. Both are able to handle a large number of sequencing reads and from portable devices (the browser-based running on a tablet), perform its task within seconds, and consume an amount of bandwidth compatible with mobile broadband networks. Such client-server approaches could develop in the future, allowing a fully automated processing of sequencing data and routine instant quality check of sequencing runs from desktop sequencers. A web access is available at http://tapir.cbs.dtu.dk. The source code for a python command-line client, a server, and supplementary data are available at http://bit.ly/1aURxkc.