The simple fool's guide to population genomics via RNA-Seq: an introduction to high-throughput sequencing data analysis

The simple fool's guide to population genomics via RNA-Seq: an introduction to high-throughput sequencing data analysis
复制标题

DOI:
10.1111/1755-0998.12003
复制
发表时间:
2012-11-01
影响因子:
7.7
通讯作者:
Palumbi, Stephen R.
Palumbi, Stephen R.
中科院分区:
生物学1区
文献类型:
--
作者:
De Wit, Pierre;Pespeni, Melissa H.;Palumbi, Stephen R.

文献摘要

被引文献

相似文献

高通量测序技术目前正在彻底改变生物学和医学领域,但分析非常大的数据集的生物信息学挑战已经减缓了人口生物学家社区对这些技术的采用。我们通过RNA-seq(SFG)介绍了简单的傻瓜指南,该文件旨在作为一个易于遵循的协议,引导用户通过非模型生物的高通量测序数据分析的一个例子。它绝不是一个详尽的协议,而是作为人口基因组学中使用的生物信息学方法的介绍,使用户能够熟悉基本的分析步骤。SFG由两部分组成。本文档总结了所需的步骤,并列出了每个步骤的基本主题和一个简单的方法。第二个文档是完整的SFG,可在http://sfg.stanford.edu上公开获得,其中包括用于数据处理和分析的详细协议,沿着定制脚本和示例文件的存储库。SFG中包括的步骤范围从组织收集到从头组装、blast注释、比对、基因表达、功能富集、SNP检测、主成分和FST离群值分析。虽然人口基因组学的技术方面正在迅速变化,但我们希望这份文件将帮助人口生物学家在高通量测序和生物信息学方面几乎没有背景,更快地采用这些新技术。
High-throughput sequencing technologies are currently revolutionizing the field of biology and medicine, yet bioinformatic challenges in analysing very large data sets have slowed the adoption of these technologies by the community of population biologists. We introduce the Simple Fool's Guide to Population Genomics via RNA-seq (SFG), a document intended to serve as an easy-to-follow protocol, walking a user through one example of high-throughput sequencing data analysis of nonmodel organisms. It is by no means an exhaustive protocol, but rather serves as an introduction to the bioinformatic methods used in population genomics, enabling a user to gain familiarity with basic analysis steps. The SFG consists of two parts. This document summarizes the steps needed and lays out the basic themes for each and a simple approach to follow. The second document is the full SFG, publicly available at http://sfg.stanford.edu, that includes detailed protocols for data processing and analysis, along with a repository of custom-made scripts and sample files. Steps included in the SFG range from tissue collection to de novo assembly, blast annotation, alignment, gene expression, functional enrichment, SNP detection, principal components and FST outlier analyses. Although the technical aspects of population genomics are changing very quickly, our hope is that this document will help population biologists with little to no background in high-throughput sequencing and bioinformatics to more quickly adopt these new techniques.