Further development of the QuickGO web interface for browsing and retrieving Gene Ontology Annotation data
Further development of the QuickGO web interface for browsing and retrieving Gene Ontology Annotation data
批准号:
BB/E023541/1
负责人:
Rolf Apweiler
金额:
$10.83万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2007
资助国家:
英国
项目状态:
已结题
起止时间:
2007 至 --
中文摘要
基因本体论(GO)联盟开发了一个以标准化格式描述基因和基因产品的本体论。GO由三个结构化词汇组成,分别描述分子功能、生物过程和细胞成分。GO已成为注释基因产品的黄金标准,因为它促进了来自同一或多个物种的基因产品的有效检索和比较。目前,在许多模式生物和基因组注释数据库中,大约有20,000个GO术语用于描述基因产物。为了支持标准化命名法,UniProt小组加入了GO注释工作,并发起了基因本体论注释(GOA)项目,为蛋白质,特别是从人类蛋白质组分配GO术语。除了手动注释外,GOA还为100,000多个物种提供了自动电子注释条目,并用GO联盟成员和同事的注释补充了GOA-UniProtKB数据。GoA是Go Consortium批注工作中最大、最全面的开源批注贡献者。围棋注释可以通过EBI或Go ftp站点下载,也可以从各种围棋浏览器查询。然而,尽管有许多围棋工具和浏览器可用,但它们都不能执行围棋用户经常要求的所有任务(围棋联盟调查,2005年10月)。没有干实验室经验的生物学家更喜欢查询简单的基于网络的界面,但可能想要检索基因列表的GO注释并链接到其他数据库,而生物信息学家更喜欢下载特定格式的大量数据。然而,通过现有接口不可能进行批量检索,并且来自不同数据库的标识符之间的映射存在问题。QuickGO是最早的基于Web的围棋浏览器之一,众所周知,并得到了广泛的使用。QuickGO最初被设计为UniProtKB策展人的注释辅助工具,当其他人发现它也有用时,QuickGO得到了进一步的发展。目前,它只允许用户搜索核心围棋数据和注释到单个UniProtKB访问、InterPro ID或酶委员会(EC)编号。随着用户请求数量的增加和对新功能的需求,它需要进一步开发,以跟上用户社区日益增长的需求。此外,Goa-UniProtKB基因关联文件正在不断增加并变得笨拙,特别是对于只对注释子集感兴趣的用户。从该文件中检索特定数据正变得单调乏味,而用于单个或批量检索GO和GOA数据的新方法是必不可少的。通过电子邮件和调查收集的用户需求表明,需要一种简单的基于网络的工具来执行对具有任何标识符的围棋注释的批量查询,并查看和下载所有或多组注释。因此,我们的目标是扩展QuickGO浏览器,以支持以下请求:-单一或批量搜索/提取GOA关联文件、FASTA或UniProtKB格式的数据,用于UniProtKB ID或登录号搜索的单个或批次基因或蛋白质-GO术语搜索/提取标注为GO术语的所有基因或蛋白质-从Unigene、DDBJ/EMBL/GenBank、Entrez、EnSembl、International Protein Index(IPI)、RefSeq等单一或批量搜索替代ID。GO本体和GOA数据已通过引用的数量和网络请求的规模以及从ftp网站下载的数据证明了它们的受欢迎程度。GO的使用支持以标准化的方式传播生物信息,我们应该提供必要的工具来支持和鼓励这些活动。通过扩展已经流行的QuickGO工具以促进更好的数据检索和操作,我们将使GO的追随者-广大科学界受益,并为普通科学家和生物信息学家提供对数据的高效访问。
英文摘要
The Gene Ontology (GO) Consortium has developed an ontology for the description of genes and gene products in a standardised format. GO consists of three structured vocabularies to describe molecular function, biological process and cellular component. GO has become the gold standard for annotating gene products as it facilitates the efficient retrieval and comparison of gene products from the same or multiple species. Currently there are ~20,000 GO terms used to describe gene products in many model organism and genome annotation databases. In support of standardized nomenclature, the UniProt group joined the GO annotation effort and initiated the Gene Ontology Annotation (GOA) project to provide assignments of GO terms to proteins, particularly from the human proteome. In addition to manual annotation, GOA also provides automated in silico annotated entries for over 100,000 species, and supplements the GOA-UniProtKB data with annotations from the GO Consortium members and associates. GOA is the largest and most comprehensive open source contributor of annotations to the GO Consortium annotation effort. GO annotation can be downloaded via EBI or GO ftp sites or queried from various GO browsers. However, despite the fact that there are many GO tools and browsers available, none of these perform all the tasks frequently requested by GO Users (GO Consortium Survey, Oct 2005). Biologists with little dry lab experience prefer to query simple web-based interfaces, but may want to retrieve GO annotations for lists of genes and link to other databases, while Bioinformaticians prefer to download bulk data in specific formats. However, batch retrieval is not possible through existing interfaces, and problems exist with mapping between identifiers from different databases. QuickGO was one of the first web-based GO browsers and is well known and used extensively. QuickGO was initially designed as an annotation aid for UniProtKB curators, and developed further when others found it useful too. Currently, it simply enables users to search core GO data and annotations to single UniProtKB accessions, InterPro IDs or Enzyme Commission (EC) numbers. With the number of user requests increasing and new functionalities required, it needs to be developed further to keep up with the ever-increasing demands of the user community. In addition, the GOA-UniProtKB gene association file is constantly increasing and becoming unwieldy, especially for users only interested in a subset of annotations. Retrieving specific data from this file is becoming tedious, and new methods for single or bulk retrieval of GO and GOA data are essential. User requirements collected via e-mails and surveys indicate a need for a simple web-based tool to perform batch queries for GO annotation with any identifier, and to view and download ALL or sets of annotations. Our objectives are therefore to extend the QuickGO browser to enable the following requests: - Single or batch search /extraction of data in either GOA association file, FASTA or UniProtKB format for single or batches of genes or proteins searched by UniProtKB ID or accession number - GO term search /extraction of all genes or proteins annotated to a GO term - Single or batch searches for alternative IDs from e.g. UniGene, DDBJ/EMBL/GenBank, Entrez, Ensembl, International Protein Index (IPI), RefSeq, etc. The GO ontologies and GOA data have proven their popularity through the number citations and the scale of web requests and data downloads from ftp sites. The use of GO supports the communication of biological information in a standardised way, and we should be providing the tools necessary to support and encourage these activities. By extending the already popular QuickGO tool to facilitate better data retrieval and manipulation, we will benefit the enormous scientific community who are followers of GO, and provide efficient accessibility to the data for bench scientists and Bioinformaticians.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1093/bioinformatics/btp536
发表时间:
2009-11-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
[Binns D, Dimmer E, Huntley R, Barrell D, O'Donovan C, Apweiler R]
通讯作者:
Apweiler R
DOI:
10.1093/nar/gkn803
发表时间:
2009-01
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Barrell D, Dimmer E, Huntley RP, Binns D, O'Donovan C, Apweiler R]
通讯作者:
Apweiler R
DOI:
10.1093/database/bap010
发表时间:
2009
期刊:
Database : the journal of biological databases and curation
影响因子:
--
作者:
[Huntley RP, Binns D, Dimmer E, Barrell D, O'Donovan C, Apweiler R]
通讯作者:
Apweiler R
ARGENT: ARgentinian GEnomics for Tuberculosis
-
批准号:EP/T015446/1
-
项目类别:Research Grant
-
资助金额:$120.62万
-
财政年份:2019
-
负责人:Rolf Apweiler
-
依托单位:
Database on demand - creating customized sequence databases for efficient protein identification
-
批准号:BB/F016255/1
-
项目类别:Research Grant
-
资助金额:$6.16万
-
财政年份:2008
-
负责人:Rolf Apweiler
-
依托单位:
Embracing new technologies to streamline improve and sustain InterPro and its contributing databases
-
批准号:BB/F010508/1
-
项目类别:Research Grant
-
资助金额:$86.45万
-
财政年份:2008
-
负责人:Rolf Apweiler
-
依托单位:
ProteomeHarvest - Excel/XML Bridge for User-friendly Proteomics Data Collection
-
批准号:BB/E00573X/1
-
项目类别:Research Grant
-
资助金额:$6.4万
-
财政年份:2006
-
负责人:Rolf Apweiler
-
依托单位:
国内基金
海外基金
登录
查看更多内容
损伤线粒体传递机制介导成纤维细胞/II型肺泡上皮细胞对话在支气管肺发育不良肺泡发育阻滞中的作用
-
批准号:82371721
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:王星云
-
依托单位:
增强子在小鼠早期胚胎细胞命运决定中的功能和调控机制研究
-
批准号:82371668
-
项目类别:面上项目
-
资助金额:52.00万元
-
批准年份:2023
-
负责人:乔云波
-
依托单位:
MAP2的m6A甲基化在七氟烷引起SST神经元树突发育异常及精细运动损伤中的作用机制研究
-
批准号:82371276
-
项目类别:面上项目
-
资助金额:47.00万元
-
批准年份:2023
-
负责人:严佳
-
依托单位:
"胚胎/生殖细胞发育特性激活”促进“神经胶质瘤恶变”的机制及其临床价值研究
-
批准号:82372327
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:马展
-
依托单位:
Irisin通过整合素调控黄河鲤肌纤维发育的分子机制研究
-
批准号:32303019
-
项目类别:青年科学基金项目
-
资助金额:30.00万元
-
批准年份:2023
-
负责人:职韶阳
-
依托单位:
TMEM30A介导的磷脂酰丝氨酸外翻促进毛细胞-SGN突触发育成熟的机制研究
-
批准号:82371172
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:杨光
-
依托单位:
HER2特异性双抗原表位识别诊疗一体化探针研制与临床前诊疗效能研究
-
批准号:82372014
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:魏伟军
-
依托单位:
水稻边界发育缺陷突变体abnormal boundary development(abd)的基因克隆与功能分析
-
批准号:32070202
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2020
-
负责人:汪泉
-
依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
-
批准号:--
-
项目类别:--
-
资助金额:40万元
-
批准年份:2020
-
负责人:Vikrant Gupta
-
依托单位:
细胞核分布基因NudCL2在细胞迁移及小鼠胚胎发育过程中的作用及机制研究
-
批准号:31701214
-
项目类别:青年科学基金项目
-
资助金额:25.0万元
-
批准年份:2017
-
负责人:张雯
-
依托单位: