Moving Just Enough Deep Sequencing Data to Get the Job Done

Moving Just Enough Deep Sequencing Data to Get the Job Done
复制标题

DOI:
10.1177/1177932219856359
复制
发表时间:
2019-06-14
影响因子:
5.8
通讯作者:
Feltus, F. Alex
Feltus, F. Alex
中科院分区:
其他
文献类型:
--
作者:
Mills, Nicholas;Bensman, Ethan M.;Feltus, F. Alex

文献摘要

被引文献

相似文献

动机:随着高通量DNA序列数据集的规模持续增长,传输和存储数据集的成本可能会阻止它们在除了最大的数据中心或商业云提供商之外的所有数据中心进行处理。为了降低这一成本,它应该是可能的,以处理只有一个子集的原始数据,同时仍然保留生物信息的interests.RESULTS:使用4个高通量的DNA序列数据集的不同测序深度从2个物种作为用例,我们证明了处理部分数据集上的数量检测到的RNA转录本使用RNA-Seq工作流的效果。我们使用转录检测来决定临界点。然后,我们物理传输最小部分数据集,并与传输完整数据集进行比较,这表明总传输时间减少了约25%。这些结果表明,随着测序数据集变得越来越大,加速分析的一种方法是简单地传输仍然足以检测生物信号的最小量的数据。可用性:所有结果均使用NCBI的公共数据集和公开可用的开源软件生成。
MOTIVATION: As the size of high-throughput DNA sequence datasets continues to grow, the cost of transferring and storing the datasets may prevent their processing in all but the largest data centers or commercial cloud providers. To lower this cost, it should be possible to process only a subset of the original data while still preserving the biological information of interest.RESULTS : Using 4 high-throughput DNA sequence datasets of differing sequencing depth from 2 species as use cases, we demonstrate the effect of processing partial datasets on the number of detected RNA transcripts using an RNA-Seq workflow. We used transcript detection to decide on a cutoff point. We then physically transferred the minimal partial dataset and compared with the transfer of the full dataset, which showed a reduction of approximately 25% in the total transfer time. These results suggest that as sequencing datasets get larger, one way to speed up analysis is to simply transfer the minimal amount of data that still sufficiently detects biological signal.AVAILABILITY: All results were generated using public datasets from NCBI and publicly available open source software.