Computational approaches to mine publicly available databases.

Computational approaches to mine publicly available databases.
复制标题

挖掘公开可用数据库的计算方法。

DOI:
10.1007/978-1-62703-980-2_24
复制
发表时间:
2014
期刊:
Methods in molecular biology (Clifton, N.J.)
影响因子:
--
通讯作者:
Berglund,JAndrew
Berglund,JAndrew
中科院分区:
--
文献类型:
--
作者:
Voelker,RodgerB;Cresko,WilliamA;Berglund,JAndrew

文献摘要

相似文献

公开可用的序列注释数据是研究人员的重要资源。可以获得许多类型的信息,包括结构注释(即,基因组特征的位置和身份)和功能注释(例如,基因表达和蛋白质相互作用)。注释数据对于询问下一代测序数据(例如,识别与作图读数相关的基因组特征)特别有用。此外,可用的海量数据为研究人员提供了挖掘现有数据集和做出新发现的机会。有效地获取、处理和询问这些数据的能力是一项有价值的、令人振奋的技能。在本章中,我们将介绍几种主要的数据存储库,并描述最常见的文件格式。为了突出使用注释数据所涉及的一些关键概念、操作和实用程序,我们提供了一个完整的使用注释来回答有关特定芯片序列数据集的基本问题的示例。
Publicly available sequence annotation data is a vital resource for researchers. Many types of information are available, including structural annotations (i.e., the locations and identities of genomic features) and functional annotations (e.g., gene expression and protein interactions). Annotation data is especially useful for interrogating Next-Gen sequencing data (e.g., identifying genomic features that are associated with mapped reads). Additionally, the vast amount of data that is available offers researchers the opportunity to mine existing data sets and make new discoveries. The ability to efficiently obtain, manipulate, and interrogate this data is a valuable and empowering skill. In this chapter, we introduce several primary data repositories and describe the most commonly encountered file formats. In order to highlight some of the key concepts, operations, and utilities that are involved in working with annotation data we provide a fully worked example of using annotations to answer some basic questions about a particular CHIP-seq data set.