Expanding the knowledge of structures and functional information through the SIFTS resource
Expanding the knowledge of structures and functional information through the SIFTS resource
批准号:
BB/M011674/1
负责人:
Sameer Velankar
金额:
$65.71万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --
中文摘要
在过去十年中,我们看到生物数据的数量和多样性迅速增加。在这一进程开始时,科学界面临的挑战是建立必要的基础设施,以有效的方式收集、管理和向研究界提供这些数据。这将生命科学研究转变为数据驱动的科学领域。但是科学界很快意识到,除了拥有这些数据之外,真正的挑战是为这些数据的生物学背景增加重要的内容,并使这些知识可供研究人员使用。对于不断增加的关于大分子三维结构的数据量来说尤其如此。大分子结构数据可以为深入了解大分子的功能机制提供重要依据。通过将其与其他生物学数据相结合,可以更好地了解生命和疾病过程,从而通过设计新的药物分子来制定更好的干预策略。大分子结构数据还可用于预测在人群中自然发现的遗传变异对大分子功能的影响,从而更好地了解遗传疾病。因此,如果我们想要利用这些数据并为不断增加的基因组和蛋白质组学信息增加价值,那么为不断增加的大分子结构数据提供生物学背景是至关重要的。SIFTS资源通过整合来自不同生物数据库的注释,将大分子结构数据(存档在两个公开可用的数据库PDB和EMDB中)链接到其生物学背景,主要是通过将其链接到UniProt, UniProt是一个公开可用的蛋白质序列数据库,处于蛋白质注释的前沿。该资源建立于2002年,多年来通过整合来自不同数据库的越来越多的蛋白质相关注释而不断发展。在SIFTS资源建立之前,每个主要的生物数据资源或研究实验室都必须建立流程和复杂的基础设施,以获得将大分子结构数据与其他数据库连接起来的必要信息。随着测序技术的快速发展,越来越多的变异和异构体信息现在变得可用。至关重要的是,SIFTS资源扩展到将这些变体和同工异构体映射到大分子结构上,并使其免费提供,以造福生命科学研究界。这将要求SIFTS资源首次更新其流程和基础设施,以包括基因组和变异信息。这些数据和相关非特征序列的扩展注释将有助于开发预测结构-功能关系的方法。这些考虑和用户要求促成了拟议的发展。拟议项目的主要目标包括-1。增强SIFTS资源中可用的注释,以包括基因组和变异信息。通过包括同工异构体、变异和相关的未表征序列来增加蛋白质序列空间的覆盖率。实现一种机制,提供特定于UniProtKB数据库中同种异构体和变异的序列注释。开发必要的基础设施,包括基于配体结合位点和组装界面残留物的增值结构注释。巩固软件流程和数据库基础设施以实现长期可持续性。
英文摘要
Over the last decade we have seen rapid increase in the amount and diversity of biological data. At the beginning of this process the challenge before the scientific community was to create the necessary infrastructure to collect, manage and make these data available in an efficient manner to the research community. This has transformed life-science research into a data driven scientific field. But very quickly the scientific community has realized that apart of having these data available, the real challenge is to add significantly to the biological context of these data and make this knowledge available to the researchers. This is especially true for the increasing amount of data on three-dimensional structures of macromolecules. The macromolecular structure data can provide great insights into the functional mechanism of the macromolecules. By integrating it with other biological data better understanding of life and disease processes can be derived leading to better intervention strategies by designing new drug molecules. The macromolecular structure data can also be used to predict the effects of genetic variation, found naturally in the population, on the function of the macromolecules again leading to better understanding of genetic diseases. So providing biological context to the increasing amount of macromolecular structure data is critical if we want to exploit these data and add value to the increasing amount of genomic and proteomic information. The SIFTS resource links the macromolecular structure data (archived in two publicly available databases PDB and EMDB) to its biological context by integrating annotations from different biological databases mainly through linking it to UniProt, a publicly available database of protein sequences, which is at the forefront of protein annotation. This resource was established in 2002 and has evolved over the years by integrating increasing number of protein related annotations from different databases. Before the SIFTS resource was established every major biological data resource or research laboratory had to establish processes and complex infrastructure to derive necessary information linking macromolecular structure data to other databases. With rapid advances in sequencing technology, an increasing amount of variation and isoform information is now becoming available. It is critical that the SIFTS resource is extended to map these variants and isoforms onto macromolecular structures and make it freely available for the benefit of the life-science research community. This will require the SIFTS resource to update its processes and infrastructure to include genomic and variation information for the first time. These data and the extended annotations for related uncharacterised sequences will be useful for developing methodologies for predicting structure-function relationship. These considerations and user requests have contributed to the proposed developments. The main objectives of the proposed project include -1. Enhance the annotations available in the SIFTS resource to include genomic and variation information.2. Increase coverage of protein sequence space by including isoforms, variants and related uncharacterised sequences.3. Implement a mechanism to provide sequence annotations specific to isoforms and variation in UniProtKB database.4. Develop the necessary infrastructure to include value-added structure-based annotations on ligand binding sites and assembly interface residues.5. Consolidate the software processes and the database infrastructure for long-term sustainability.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.3390/metabo11010048
发表时间:
2021-01-12
期刊:
Metabolites
影响因子:
4.1
作者:
[Feuermann M, Boutet E, Morgat A, Axelsen KB, Bansal P, Bolleman J, de Castro E, Coudert E, Gasteiger E, Géhant S, Lieberherr D, Lombardot T, Neto TB, Pedruzzi I, Poux S, Pozzato M, Redaschi N, Bridge A, On Behalf Of The UniProt Consortium]
通讯作者:
On Behalf Of The UniProt Consortium
DOI:
10.1093/nar/gkaa1113
发表时间:
2021-01-08
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Gene Ontology Consortium]
通讯作者:
Gene Ontology Consortium
DOI:
10.1016/j.sbi.2016.06.018
发表时间:
2016-10
期刊:
CURRENT OPINION IN STRUCTURAL BIOLOGY
影响因子:
6.8
作者:
[Berman, Helen M., Burley, Stephen K., Kleywegt, Gerard J., Markley, John L., Nakamura, Haruki, Velankar, Sameer]
通讯作者:
Velankar, Sameer
DOI:
10.1093/nar/gkx1070
发表时间:
2018-01-04
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Mir S, Alhroub Y, Anyango S, Armstrong DR, Berrisford JM, Clark AR, Conroy MJ, Dana JM, Deshpande M, Gupta D, Gutmanas A, Haslam P, Mak L, Mukhopadhyay A, Nadzirin N, Paysan-Lafosse T, Sehnal D, Sen S, Smart OS, Varadi M, Kleywegt GJ, Velankar S]
通讯作者:
Velankar S
BBSRC-NSF/BIO: An AI-based domain classification platform for 200 million 3D-models of proteins to reveal protein evolution
-
批准号:BB/Y000455/1
-
项目类别:Research Grant
-
资助金额:$46.4万
-
财政年份:2024
-
负责人:Sameer Velankar
-
依托单位:
20-BBSRC/NSF-BIO: From atoms to molecules to cells - Multi-scale tools and infrastructure for visualization of annotated 3D structure data
-
批准号:BB/W017970/1
-
项目类别:Research Grant
-
资助金额:$56.55万
-
财政年份:2023
-
负责人:Sameer Velankar
-
依托单位:
FUNCLAN - FUNctional annotations through Conformational Landscape Analysis
-
批准号:BB/V016113/1
-
项目类别:Research Grant
-
资助金额:$51.33万
-
财政年份:2022
-
负责人:Sameer Velankar
-
依托单位:
CIBR 19-BBSRC-NSF/BIO: Next generation PDB - FACT infrastructure with value added FAIR data supporting diverse research and education user communities
-
批准号:BB/V004247/1
-
项目类别:Research Grant
-
资助金额:$48.28万
-
财政年份:2021
-
负责人:Sameer Velankar
-
依托单位:
BioChemGRAPH - an integrated knowledge graph to facilitate basic and translational research
-
批准号:BB/T01959X/1
-
项目类别:Research Grant
-
资助金额:$69.31万
-
财政年份:2020
-
负责人:Sameer Velankar
-
依托单位:
Increasing the Coverage and Accuracy of CATH for Comparative Genomics and Variant Interpretation
-
批准号:BB/R015201/1
-
项目类别:Research Grant
-
资助金额:$13.41万
-
财政年份:2019
-
负责人:Sameer Velankar
-
依托单位:
3D-Gateway to protein structure and function
-
批准号:BB/S020071/1
-
项目类别:Research Grant
-
资助金额:$61.36万
-
财政年份:2019
-
负责人:Sameer Velankar
-
依托单位:
BBSRC-NSF/BIO - Expanding fold library in the twilight zone to facilitate structure determination of macromolecular machines
-
批准号:BB/S017135/1
-
项目类别:Research Grant
-
资助金额:$43.0万
-
财政年份:2019
-
负责人:Sameer Velankar
-
依托单位:
FunPDBe - enhancing structural and functional annotation of macromolecular structure data in the PDB by collaboration and integration
-
批准号:BB/P024351/1
-
项目类别:Research Grant
-
资助金额:$73.97万
-
财政年份:2017
-
负责人:Sameer Velankar
-
依托单位:
India partnering award: Sustainable data archiving and dissemination strategy to support data driven biology
-
批准号:BB/P025846/1
-
项目类别:Research Grant
-
资助金额:$3.89万
-
财政年份:2017
-
负责人:Sameer Velankar
-
依托单位:
PDBHarvest - Harvesting more and better metadata from CCP4 projects to enrich structure depositions to the PDB
-
批准号:BB/M020428/1
-
项目类别:Research Grant
-
资助金额:$8.94万
-
财政年份:2015
-
负责人:Sameer Velankar
-
依托单位:
BioSolr: addressing the challenges in making biomedical data easily accesible using the world-leading Apache-Solr search-engine framework
-
批准号:BB/M013146/1
-
项目类别:Research Grant
-
资助金额:$19.12万
-
财政年份:2014
-
负责人:Sameer Velankar
-
依托单位:
海外基金