课题基金 / 基金详情

ABI Innovation: Identification and correction of incorrect sequences in reference protein databases

ABI Innovation: Identification and correction of incorrect sequences in reference protein databases
ABI Innovation:参考蛋白质数据库中错误序列的识别和纠正
批准号:
1759625
负责人:
William Pearson
金额:
$82.58万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-05-01 至 2022-06-30

项目摘要

项目成果

William Pearson的其他基金

相似基金

相关文献

中文摘要
翻译
随着人类基因组测序和基因组时代的到来,医生和科学家们有了一份完整的“部件清单”——DNA和蛋白质序列——从人类到植物再到细菌,数千种不同生物细胞中的机器。但是,因为我们不能直接看到机器或它们的部件,所以一些细节是不正确的。这些方法产生了不正确的细节,部分原因是它们是在“前基因组”时代发展起来的,当时我们的部件列表远不完整,部分原因是我们使用来自一种生物体(小鼠)的部件列表来帮助我们识别另一种生物体(大鼠)的部件。但是,如果第一组部分(蛋白质序列)有错误,这些错误可能会出现在第二组中,随着对更多生物体的研究,错误会成倍增加。不正确的蛋白质序列使得识别可能与癌症和其他疾病相关的突变变得更加困难。今天,蛋白质序列的质量控制很大程度上依赖于其他蛋白质序列的准确性。该提案旨在极大地扩展用于确认细胞机制细节(序列)正确的信息类型,并与Uniprot蛋白质序列数据库合作,以确保错误的序列得到纠正,从而使未来的分析不再重复错误。拟议中的研究将增加科学家的信心,使他们相信表型之间令人兴奋的差异,例如患病细胞和正常细胞,或不同的粮食作物品系,是真实的,而不是错误的蛋白质序列的结果。UniProt参考蛋白集是一个重要的资源,为基因组生物学、结构/功能预测和个性化医疗的进步提供了基础。但参考蛋白集存在特征误差,在基因组标注中会被放大,降低了功能标注的准确性和相似性搜索的灵敏度。该项目是与Uniprot蛋白质资源的合作,将开发识别蛋白质序列错误的新策略;这些策略将被合并到Uniprot注释管道中,以从参考数据库中删除错误序列。最初,错误将通过将外显子边界和Pfam结构域注释整合到基于blast的相似性搜索结果中来识别当前识别的错误,例如缺失/额外的外显子和部分结构域。此外,还将制定策略,通过聚集相似的蛋白质并寻找“异常值”来检测目前未被识别的错误。“可疑”蛋白质序列列表将由Uniprot注释器检查并分类为正确或错误的错误,成功的错误检测策略将被改进并集成到Uniprot注释管道中,从而减少功能性注释错误。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
With the sequencing of the Human genome and the dawn of the genome age, physicians and scientists have a complete 'parts-list' - DNA and protein sequences - for the machinery in the cells of thousands of different organisms, from people to plants to bacteria. But, because we cannot look at the machines or their parts directly, some of the details are incorrect. The methods produce incorrect details in part because they were developed in the 'pre-genome' era, when our parts lists were far less complete, and in part because we use the parts-list from one organism (mouse) to help us identify the parts in another organism (rat). But, if the first set of parts (protein sequences) has errors, those errors can end up in the second set, and as more organisms are studied, the errors multiply. Incorrect protein sequences make it more difficult to identify mutations that may be associated with cancer and other diseases. Today, protein sequence quality control largely relies on the accuracy of other protein sequences. This proposal seeks to greatly expand the types of information used to confirm that the details (sequences) of the cellular machinery are correct, and to collaborate with the Uniprot protein sequence database to ensure that incorrect sequences are corrected and so that future analyses do not repeat mistakes. The proposed research will increase researchers' confidence that exciting differences between phenotypes, such as diseased and normal cells, or different strains of food crops, are genuine, and not the result of mistaken protein sequences.The UniProt Reference protein sets are a critical resource, providing a foundation for advances in genome biology, structure/function prediction, and personalized medicine. But Reference protein sets have characteristic errors that can be amplified in genome annotation, reducing the accuracy of functional annotation and similarity search sensitivity. This project is a collaboration with the Uniprot protein resource that will develop novel strategies for identifying protein sequence errors; these strategies will be incorporated into the Uniprot annotation pipelines to remove erroneous sequences from Reference databases. Initially, errors will be found by integrating exon-boundary and Pfam domain annotations into BLAST-based similarity search results to identify currently recognized errors, such as missing/additional exons and partial domains. In addition, strategies will be developed to detect currently unrecognized errors, by clustering similar proteins and looking for 'outliers'. Lists of 'suspect' protein sequences will be examined by Uniprot annotators and classified as true or mistaken errors, and successful error detection strategies will be refined and integrated into the Uniprot annotation pipeline, reducing functional annotation errors.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Urban Cities' Self-Study Grant
海外基金