PIR Classification Database for Genomic Research
PIR Classification Database for Genomic Research
批准号:
9974855
负责人:
Cathy Wu
金额:
$28.31万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1999
资助国家:
美国
项目状态:
已结题
起止时间:
1999-10-01 至 2002-09-30
中文摘要
DBI 9974855;吴,凯西。国家生物医学研究基金会。PIR基因组研究分类数据库项目概述随着人类基因组计划和其他类似的大型测序计划,分子序列数据继续呈指数级增长,充分了解基因组已成为计算分子生物学的一个巨大挑战。需要先进的数据库,以便于从海量数据中检索相关信息,并提供对蛋白质结构和功能的洞察。这项研究的主要目标是开发一个原型关系分类数据库,为描述蛋白质的全面家族关系和结构和功能特征提供一个集成的平台。该数据库将结合几个广泛使用的序列、结构和功能数据库的分类方案,允许用户检索任何给定序列的方便显示的摘要,并查询具有选定属性的序列和家族。数据库本身将从源数据库自动生成并频繁更新。对于所有蛋白质的完全和非重叠放置,将使用PIR超家族/家族分类作为基本组织方案。基本超家族实体将具有由其他分类方案定义的属性,例如成员资格、注释以及与家族、域和主题的关系。此外,还将汇编一个补充主题集,为其他主题数据库中找不到的功能和结构主题提供灵活的诊断模式。拟议的研究将建立在我们现有数据库的基础上,并在此基础上扩展。Oracle关系实现将支持数据库查询,以询问有关蛋白质序列或家族的属性组合及其关系的问题。用户将能够执行序列搜索,对数据库进行查询,并提供反馈信息。该报告将包括特征和关系的图形显示,以及到底层数据库的超文本链接。拟议的分类数据库的框架有几个特点。PIR超家族是现有的唯一支持所有蛋白质包容和唯一聚类的分类方案,这是完整数据库组织的先决条件。该设计支持基于全局和基序序列相似性的基因组注释,以及序列注释。分类数据库中的每个超家族记录将比任何其他单一信息资源更全面地记录家族关系和特征。该数据库将是研究人员的一个非常有用的工具,也是建立一个更广泛的数据库的原型,该数据库包括更多类型的信息以及与其他结构和功能分类数据库的连接。分类数据库将是一个重要的信息资源,有助于识别新的基因组序列,推断新的生物学知识,并提高数据库的完整性。
英文摘要
DBI 9974855; Wu, Cathy. National Biomedical Research Foundation. PIR Classification Database for Genomic ResearchProject SummaryAs molecular sequence data continue to grow exponentially due to the Human Genome Project and other similar large sequencing projects, gaining a full understanding of the genome has become a great challenge in computational molecular biology. Advanced databases are needed to facilitate the retrieval of relevant information from the voluminous data and to provide insight into protein structure and function. The major objective of this research is to develop a prototype relational classification database that will provide an integrated platform for describing comprehensive family relationships and structural and functional features of proteins. The database will combine the classification schemes of several widely used sequence, structure, and function databases, allowing users to retrieve conveniently displayed summaries for any given sequence and to query for sequences and families with selected properties. The database itself will be automatically generated from the source databases and updated frequently. For complete and non-overlapping placement of all proteins, the PIR superfamily/family classification will be used as the underlying organization scheme. The basic superfamily entity will have attributes such as membership, annotation, and relationship to families, domains, and motifs defined by other classification schemes. In addition, a supplementary motif collection will be compiled to provide flexible and diagnostic patterns for functional and structural motifs not found in other motif databases. The proposed research will build upon and extend from our current databases. The ORACLE relational implementation will support database queries for asking questions concerning combinations of attributes for protein sequences or families and their relationships. Users will be able to perform sequence searches, query against the databases, and offer feedback information. The report will include graphical displays of features and relationships, as well as hypertext links to underlying databases. The framework of the proposed classification database has several features. The PIR superfamily is the only existing classification scheme supporting inclusive and unique clustering of all proteins, which is a prerequisite for complete database organization. The design supports genomic annotation based on global and motif sequence similarities, as well as sequence annotation. Each superfamily record in the classification database will document family relationships and features more comprehensively than any other single information resource. This database will be both a very useful tool for researchers and a prototype for construction of a more extensive database that includes more types of information and connections to additional structure and function classification databases. The classification database will be a significant informational resource to assist the identification of new genomic sequences, the inference of new biological knowledge, and the improvement of database integrity.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Learning for Generalization in Large-Scale Cyber-Physical Systems
-
批准号:2239566
-
项目类别:Continuing Grant
-
资助金额:$55.0万
-
财政年份:2023
-
负责人:Cathy Wu
-
依托单位:
Collaborative Research: CPS: Medium: An Online Learning Framework for Socially Emerging Mixed Mobility
-
批准号:2149548
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2022
-
负责人:Cathy Wu
-
依托单位:
BIBM-2012 Travel Awards: Broadening Interdisciplinary Research and Education in Bioinformatics and Biomedicine- to be held in Philadelphia, PA, October 4 - 7, 2012
-
批准号:1242809
-
项目类别:Standard Grant
-
资助金额:$1.6万
-
财政年份:2012
-
负责人:Cathy Wu
-
依托单位:
ABI Development: Integrative Bioinformatics for Knowledge Discovery of PTM Networks
-
批准号:1062520
-
项目类别:Continuing Grant
-
资助金额:$159.26万
-
财政年份:2011
-
负责人:Cathy Wu
-
依托单位:
BIBM Conference: Fostering Interdisciplinary Research and Education in Bioinformatics and Biomedicine
-
批准号:0960601
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:2009
-
负责人:Cathy Wu
-
依托单位:
Linking Text Mining with Ontology and Systems Biology
-
批准号:0850319
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2009
-
负责人:Cathy Wu
-
依托单位:
Integrated Protein Classification Database System For Genomic and Proteomic Research
-
批准号:0138188
-
项目类别:Continuing Grant
-
资助金额:$49.99万
-
财政年份:2002
-
负责人:Cathy Wu
-
依托单位:
Database of Protein Modifications: Enhancements for Visualization, Modeling and Internet Access
-
批准号:9808414
-
项目类别:Standard Grant
-
资助金额:$22.04万
-
财政年份:1998
-
负责人:Cathy Wu
-
依托单位:
海外基金