课题基金 / 基金详情

项目摘要

项目成果

Alex Bateman的其他基金

相似基金

相关文献

中文摘要
翻译
项目总结 用于人工智能/机器学习的UniProt蛋白质序列和功能嵌入 准备就绪-补充请求2022 人工智能和机器学习(AI/ML)具有推动生物医学研究的潜力。这个 UniProt家长赠款(U24HG007822)的本附录申请的总体目标是:(I)支持 AI/ML社区通过提供蛋白质序列嵌入并利用这些嵌入 UniRef生产中准确快速的蛋白质聚类(II)探索UniProt的嵌入方法 功能注释数据,以及(Iii)与AI/ML社区接触,以促进UniProt的AI/ML准备工作 数据。UniProt自成立以来一直在提供蛋白质序列和注释数据方面处于领先地位 在2002年。UniProt为生物医学领域的数百个AI/ML应用程序提供黄金标准培训数据 研究。蛋白质序列嵌入展示了蛋白质聚类和结构和 功能分析和预测。通过提供UniProt蛋白质序列嵌入,我们将增加 序列嵌入的可访问性,减少社区中的重复工作,并建立标准 这有助于对模型进行评估和比较。我们将测试不同的嵌入方法,重点是 关于我们最近对用户社区的调查所确定的最广泛采用的方法。我们还将 研究使用序列嵌入来加速UniRef生产中的序列聚类。同样, UniProt中有许多功能注释,可以嵌入到AI/ML模型中使用。 作为一个测试用例,我们将探索将Rhea生化反应注释嵌入酶和 传送者。除了传播嵌入,我们还将开发可视化它们的方法,并 将它们与现有的酶分类系统进行比较。最后,为了确保我们的工作与 为满足市民的需要,我们会举办工作坊,与市民共同研究他们的用例和 嵌入的应用程序。我们将邀请各种利益相关者,包括参与 嵌入调查,金属结合位点预测挑战的参与者目前正在进行中 以及NIH的代表。总的来说,这项建议中的工作将加强 准备将UniProt用于AI/ML,并将把AI/ML方法整合到UniProt生产中。
英文摘要
PROJECT SUMMARY UNIPROT PROTEIN SEQUENCE AND FUNCTION EMBEDDINGS FOR AI/MACHINE LEARNING READINESS - SUPPLEMENT REQUEST 2022 Artificial intelligence and machine learning (AI/ML) has the potential to advance biomedical research. The overall goals of this Supplement application for the UniProt parent grant (U24HG007822) are to (i) support the AI/ML community by providing protein sequence embeddings and make use of these embeddings for accurate and fast protein clustering in UniRef production, (ii) explore methods of embedding UniProt functional annotation data, and (iii) engage with the AI/ML community to advance AI/ML readiness of UniProt data. UniProt has been a leader in the provision of protein sequence and annotation data since its inception in 2002. UniProt provides gold standard training data for hundreds of AI/ML applications in biomedical research. Protein sequence embeddings show enormous promise for protein clustering and structural and functional analysis and prediction. By providing UniProt protein sequence embeddings, we will increase the accessibility of sequence embeddings, reduce duplication of effort in the community, and establish a standard that can facilitate evaluation and comparison of models. We will test different embedding methods, focusing on the most widely adopted methods as determined by our recent survey of the user community. We will also investigate using sequence embeddings to speed up sequence clustering in UniRef production. Similarly, there are many functional annotations in UniProt that are amenable to embedding for use in AI/ML models. As a test case, we will explore embedding of Rhea biochemical reaction annotations for enzymes and transporters. In addition to disseminating the embeddings, we will develop methods to visualize them and compare them to existing enzyme classification systems. Finally, to ensure that our work aligns with community needs, we will organise a workshop to work with the community on their use cases and applications for embeddings. We will invite various stakeholders, including researchers that participated in the embeddings survey, participants of the metal binding site prediction challenge currently underway as part of the parent grant, as well as NIH representatives. Overall, the work in this proposal will enhance the readiness of UniProt for use in AI/ML and will integrate AI/ML methods into UniProt production.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
UniProt: A centralized protein sequence and function resource
UniProt: A Protein Sequence and Function Resource for Biomedical Science
UniProt - Enhancing functional genomics data access for the Alzheimer's Disease (AD) and dementia-related protein research communities
UniProt: A centralized protein sequence and function resource
海外基金