Using galaxy to perform large-scale interactive data analyses.

Using galaxy to perform large-scale interactive data analyses.
复制标题

DOI:
10.1002/0471250953.bi1005s19
复制
发表时间:
2007-09-01
影响因子:
--
通讯作者:
Nekrutenko, Anton
Nekrutenko, Anton
中科院分区:
其他
文献类型:
--
作者:
Taylor, James;Schenck, Ian;Nekrutenko, Anton

文献摘要

被引文献

相似文献

虽然大多数实验生物学家知道从哪里下载基因组数据,但很少有人有具体的计划来分析它。这种情况可以通过以下方式来纠正:(1)提供统一的门户网站来服务基因组数据;(2)构建Web应用程序来允许灵活的检索和对数据的实时分析。强大的资源,如UCSC基因组浏览器已经解决了第一个问题。然而,第二个问题仍然悬而未决。例如,如何从所有测序的哺乳动物中找到单核苷酸多态性(SNP)密度最高的人类蛋白质编码外显子,并提取正向序列?事实上,人们可以从UCSC基因组浏览器访问所有相关数据。但是,一旦数据被下载,人们将如何处理数百万个SNP和千兆字节的比对?Galaxy(http:g2.bx.psu.edu)就是专门为此目的设计的。它通过允许用户以前所未有的方式在单个界面中访问和分析数据,放大了现有资源(如UCSC基因组浏览器)的优势。
While most experimental biologists know where to download genomic data, few have a concrete plan on how to analyze it. This situation can be corrected by: (1) providing unified portals serving genomic data and (2) building Web applications to allow flexible retrieval and on-the-fly analyses of the data. Powerful resources, such as the UCSC Genome Browser already address the first issue. The second issue, however, remains open. For example, how to find human protein-coding exons with the highest density of single nucleotide polymorphisms (SNPs) and extract orthologous sequences from all sequenced mammals? Indeed, one can access all relevant data from the UCSC Genome Browser. But once the data is downloaded how would one deal with millions of SNPs and gigabytes of alignments? Galaxy (http://g2.bx.psu.edu) is designed specifically for that purpose. It amplifies the strengths of existing resources (such as UCSC Genome Browser) by allowing the user to access and, most importantly, analyze data within a single interface in an unprecedented number of ways.