Tximeta: Reference sequence checksums for provenance identification in RNA-seq

Tximeta: Reference sequence checksums for provenance identification in RNA-seq
复制标题

DOI:
10.1371/journal.pcbi.1007664
复制
发表时间:
2020-02-01
影响因子:
4.3
通讯作者:
Patro, Rob
Patro, Rob
中科院分区:
生物学2区
文献类型:
--
作者:
Love, Michael I.;Soneson, Charlotte;Patro, Rob

文献摘要

被引文献

相似文献

正确的注释元数据对于可重复和准确的RNA-seq分析至关重要。当文件被公开共享或在具有不正确或缺失注释元数据的合作者之间共享时,从原始数据再现生物信息学分析变得困难或不可能。这也使得在其适当的基因组背景下定位转录组特征(例如转录物或基因)变得更加困难,这对于将表达数据与其他数据集重叠是必要的。我们以R/Bioconductor包tximeta的形式提供了一个解决方案,该解决方案在转录量化文件的导入过程中代表用户自动执行许多注释和元数据收集任务。通过存储在定量输出中的散列校验和识别正确的参考转录组,并下载和本地缓存关键转录数据库。基于参考序列校验和自动添加注释元数据的计算范例可以通过帮助减少生物信息学分析期间的开销、防止代价高昂的生物信息学错误以及促进计算再现性来极大地促进基因组工作流。tximeta包可在https://bioconductor.org/packages/tximeta上获得。
Correct annotation metadata is critical for reproducible and accurate RNA-seq analysis. When files are shared publicly or among collaborators with incorrect or missing annotation metadata, it becomes difficult or impossible to reproduce bioinformatic analyses from raw data. It also makes it more difficult to locate the transcriptomic features, such as transcripts or genes, in their proper genomic context, which is necessary for overlapping expression data with other datasets. We provide a solution in the form of an R/Bioconductor package tximeta that performs numerous annotation and metadata gathering tasks automatically on behalf of users during the import of transcript quantification files. The correct reference transcriptome is identified via a hashed checksum stored in the quantification output, and key transcript databases are downloaded and cached locally. The computational paradigm of automatically adding annotation metadata based on reference sequence checksums can greatly facilitate genomic workflows, by helping to reduce overhead during bioinformatic analyses, preventing costly bioinformatic mistakes, and promoting computational reproducibility. The tximeta package is available at https://bioconductor.org/packages/tximeta.