Exploiting data driven computational approaches for understanding protein structure and function in InterPro and Pfam
Exploiting data driven computational approaches for understanding protein structure and function in InterPro and Pfam
批准号:
BB/S020381/1
负责人:
Alex Bateman
金额:
$103.95万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
蛋白质是具有多种关键功能的生物大分子,从酶(如负责发酵的物质)到转运体(如血液中的血红蛋白)再到机械结构(如肌肉中的肌动蛋白和肌球蛋白)。蛋白质是由称为氨基酸的线性聚合物组成的。它们通常折叠成复杂的三维(3D)结构,通常与其他蛋白质和分子相互作用以发挥其功能。蛋白质序列的知识可以促进对迄今为止未被发现的酶的了解,这些酶在生物技术领域或制药行业感兴趣的新药中具有潜在的应用价值。详细了解蛋白质的功能结构,包括氨基酸在3D结构中的排列,使科学家能够诊断疾病以及设计更有效的酶。如今,我们基于现代高通量DNA测序(HTS)技术生成新蛋白质序列的能力远远超过了我们对它们进行功能表征的能力。因此,通过识别新序列与少数实验特征示例之间的相似性,使用这些来推断功能(即注释),大多数序列都进行了计算注释。最近,高温超导技术已被直接应用于环境样本,以发现以前未培养的细菌和单细胞真核生物,并使大型和复杂的基因组得以重建,如植物。这些方法正在纠正蛋白质序列数据库中的许多历史偏差。然而,为了让人类理解和利用这些数据,需要对序列进行功能注释,这最好是使用从相关序列集(称为蛋白质家族)收集的信息来完成。InterPro是世界领先的蛋白质家族资源,它整合了来自13个不同专业数据库的信息,为用户提供全面的序列功能分析。它的成员数据库之一Pfam是包含功能注释的蛋白质结构域家族的集合。InterPro和Pfam都是蛋白质研究领域公认的主要资源。在本应用中,我们对这两种资源提出了关键的发展建议,以增强它们的实用性、功能性和可扩展性,并使它们处于独特的位置,以应对该领域迫在眉睫的进展。我们将利用与其他蛋白质数据库预先建立的链接,同时建立额外的管道,在这些现有资源和新资源之间开发和交换最新信息。我们将通过建立新的相关蛋白质集(或簇)的家族来提高来自环境来源的蛋白质序列的覆盖率。考虑到蛋白质结构和功能之间的基本关联,我们将开发一个管道,不仅可以导入Pfam条目的结构模型并通过网站呈现,还可以确保模型保持最新状态。为了增加这两种资源的覆盖范围和功能注释,我们将整合新的资源来提供子领域分类,并通过结合文献搜索和增强的管理工具来改进注释。为了完善注释,我们将在InterProScan(我们的自动蛋白质序列注释软件包)中采用一种名为TreeGrafter的新算法,并将来自PANTHER等数据库的蛋白质属性受控词汇表与InterPro中已有的词汇表集成在一起。我们将评估HMMER软件的升级版本的性能,该软件广泛用于构建蛋白质家族,包括Pfam,以提高未来的可扩展性。最后,我们将重点关注8个具有农业重要性的基因组,包括鸡、鲑鱼和小麦,通过系统地注释Pfam和扩展的InterPro中的2000个相关条目。
英文摘要
Proteins are biological macromolecules that perform a diverse array of crucial functions, from enzymes (e.g. the entities responsible for fermentation) to transporters (e.g. hemoglobin in the blood) to mechanical structures (e.g. actin and myosin in muscle). Proteins are synthesized as linear polymers of building blocks called amino acids. They usually fold into complex three-dimensional (3D) structures, and typically interact with other proteins and molecules to perform their function. Knowledge of protein sequences can facilitate insights into hitherto undiscovered enzymes with potential applications in the biotechnology sector, or novel drugs of interest to the pharmaceutical industry. Detailed understanding of the functional architecture of proteins, including the arrangement of amino acids in a 3D structure, enables scientists to diagnose diseases as well as design more effective enzymes. These days, our ability to generate new protein sequences based on modern high-throughput DNA sequencing (HTS) techniques far outstrips our ability to functionally characterise them. Thus, most sequences are computationally annotated, by identifying similarities between new sequences and the few experimentally characterised examples, using these to infer function (i.e. annotate). More recently, HTS has been applied directly to environmental samples to discover previously uncultured bacteria and single cell eukaryotes, and to enable the reconstruction of large and complex genomes, like plants. Such approaches are correcting many of the historical biases in the protein sequence databases. However, for humankind to understand and utilise these data, sequences need to be functionally annotated, which is best accomplished using the information gleaned from sets of related sequences (known as protein families). InterPro is a world leading protein family resource that merges information from 13 different specialist databases to present the user with comprehensive functional analysis of sequences. One of its member databases, Pfam, is a collection of protein domain families containing functional annotations. Both InterPro and Pfam are well-established primary resources in the field of protein research. In this application, we propose crucial developments to both of these resources in order to augment their utility, functionality and scalability, as well as uniquely position them to tackle imminent advances in the field. We will leverage pre-established links with other protein databases and concurrently build additional pipelines to develop and exchange the latest information between these existing and new resources.We will improve coverage of protein sequences originating from environmental sources by building families for novel sets (or clusters) of related proteins. Considering the fundamental association between protein structure and function, we will develop a pipeline that will not only import structural models for Pfam entries and present them via the website, but will also ensure that the models remain up to date. To increase coverage and functional annotations in both resources, we will integrate new resources to provide sub-domain classifications, and improve annotations through combined literature searches and enhanced curation tools. To refine annotations, we will adopt a new algorithm called TreeGrafter to InterProScan (our software package that performs automatic annotations of protein sequences), and integrate controlled vocabularies for protein attributes from databases like PANTHER with those already in InterPro. We will evaluate the performance of an upgraded version of the HMMER software that is widely used to build protein families, including Pfam, to improve future scalability. Finally, we will focus on eight genomes of agricultural importance, including chicken, salmon, and wheat, by systematically annotating 2000 associated entries in Pfam and by extension, InterPro.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1093/nar/gkac1098
发表时间:
2023-01-06
期刊:
NUCLEIC ACIDS RESEARCH
影响因子:
14.9
作者:
[Thakur, Matthew, Bateman, Alex, Brooksbank, Cath, Freeberg, Mallory, Harrison, Melissa, Hartley, Matthew, Keane, Thomas, Kleywegt, Gerard, Leach, Andrew, Levchenko, Mariia, Morgan, Sarah, McDonagh, Ellen M., Orchard, Sandra, Papatheodorou, Irene, Velankar, Sameer, Vizcaino, Juan Antonio, Witham, Rick, Zdrazil, Barbara, McEntyre, Johanna]
通讯作者:
McEntyre, Johanna
DOI:
10.1093/nar/gkaa977
发表时间:
2021-01-08
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Blum M, Chang HY, Chuguransky S, Grego T, Kandasaamy S, Mitchell A, Nuka G, Paysan-Lafosse T, Qureshi M, Raj S, Richardson L, Salazar GA, Williams L, Bork P, Bridge A, Gough J, Haft DH, Letunic I, Marchler-Bauer A, Mi H, Natale DA, Necci M, Orengo CA, Pandurangan AP, Rivoire C, Sigrist CJA, Sillitoe I, Thanki N, Thomas PD, Tosatto SCE, Wu CH, Bateman A, Finn RD]
通讯作者:
Finn RD
DOI:
10.1093/nar/gkaa913
发表时间:
2021-01-08
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, Tosatto SCE, Paladin L, Raj S, Richardson LJ, Finn RD, Bateman A]
通讯作者:
Bateman A
DOI:
10.1093/nar/gkab1127
发表时间:
2022-01-07
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Cantelli G, Bateman A, Brooksbank C, Petrov AI, Malik-Sheriff RS, Ide-Smith M, Hermjakob H, Flicek P, Apweiler R, Birney E, McEntyre J]
通讯作者:
McEntyre J
DOI:
10.1093/bioadv/vbac072
发表时间:
2022
期刊:
Bioinformatics advances
影响因子:
--
作者:
[]
通讯作者:
Improving accuracy, coverage, and sustainability of functional protein annotation in InterPro, Pfam and FunFam using Deep Learning methods
-
批准号:BB/X018660/1
-
项目类别:Research Grant
-
资助金额:$95.75万
-
财政年份:2024
-
负责人:Alex Bateman
-
依托单位:
UKRI/BBSRC-NSF/BIO: Unifying Pfam protein sequence and ECOD structural classifications with structure models
-
批准号:BB/X012492/1
-
项目类别:Research Grant
-
资助金额:$92.15万
-
财政年份:2023
-
负责人:Alex Bateman
-
依托单位:
Rfam: The community resource for RNA families
-
批准号:BB/S020462/1
-
项目类别:Research Grant
-
资助金额:$64.88万
-
财政年份:2019
-
负责人:Alex Bateman
-
依托单位:
RNAcentral, the RNA sequence database
-
批准号:BB/N019199/1
-
项目类别:Research Grant
-
资助金额:$87.33万
-
财政年份:2017
-
负责人:Alex Bateman
-
依托单位:
Rfam: Towards a sustainable resource for understanding the genomic functional ncRNA repertoire
-
批准号:BB/M011690/1
-
项目类别:Research Grant
-
资助金额:$54.53万
-
财政年份:2015
-
负责人:Alex Bateman
-
依托单位:
Keeping pace with protein sequence annotation; consolidating and enhancing Pfam and InterPro's methodologies for functional prediction
-
批准号:BB/L024136/1
-
项目类别:Research Grant
-
资助金额:$69.49万
-
财政年份:2014
-
负责人:Alex Bateman
-
依托单位:
The RNAcentral database of non-coding RNAs
-
批准号:BB/J019232/1
-
项目类别:Research Grant
-
资助金额:$12.67万
-
财政年份:2012
-
负责人:Alex Bateman
-
依托单位:
Embracing new technologies to streamline improve and sustain InterPro and its contributing databases
-
批准号:BB/F010435/1
-
项目类别:Research Grant
-
资助金额:$39.16万
-
财政年份:2008
-
负责人:Alex Bateman
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
复杂数据下半参数转换模型及其在老年慢性病发展中的应用研究
-
批准号:72101261
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:孙韬
-
依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
-
批准号:--
-
项目类别:--
-
资助金额:40万元
-
批准年份:2020
-
负责人:Vikrant Gupta
-
依托单位:
基于高频信息下高维波动率矩阵估计及应用
-
批准号:71901118
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2019
-
负责人:穆燕
-
依托单位:
半参数空间自回归面板模型的有效估计与应用研究
-
批准号:71961011
-
项目类别:地区科学基金项目
-
资助金额:16.0万元
-
批准年份:2019
-
负责人:丁飞鹏
-
依托单位:
高频数据波动率统计推断、预测与应用
-
批准号:71971118
-
项目类别:面上项目
-
资助金额:50.0万元
-
批准年份:2019
-
负责人:孔新兵
-
依托单位:
经济管理中复杂数据和复杂行为的分析方法及其应用
-
批准号:71931004
-
项目类别:重点项目
-
资助金额:230.0万元
-
批准年份:2019
-
负责人:周勇
-
依托单位:
基于个体分析的投影式非线性非负张量分解在高维非结构化数据模式分析中的研究
-
批准号:61502059
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2015
-
负责人:刘昶
-
依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
-
批准号:61373035
-
项目类别:面上项目
-
资助金额:77.0万元
-
批准年份:2013
-
负责人:冯志勇
-
依托单位: