课题基金 / 基金详情

项目摘要

项目成果

Catherine E. Costello的其他基金

相似基金

相关文献

中文摘要
翻译
这个子项目是许多利用资源的研究子项目之一 由NIH/NCRR资助的中心拨款提供。子项目的主要支持 而子项目的主要调查员可能是由其他来源提供的, 包括其它NIH来源。 列出的子项目总成本可能 代表子项目使用的中心基础设施的估计数量, 而不是由NCRR赠款提供给子项目或子项目工作人员的直接资金。 随着大量不同的MS仪器和数据分析软件平台可用于MS和蛋白质组学,它变得难以操纵和管理各种数据集。我们正在开发新的软件来处理蛋白质和肽MS和MS/MS数据,定位和分配翻译后修饰(PTM),并帮助将结果置于生物学背景中。BUPID程序是在Linux下用C语言开发的,通过基于CGI的Web界面可访问主程序。 编写外壳数据转换程序以实现用户友好的GUI界面,该界面可以在无人值守的批处理模式下操作。对内部获得的现有MALDI-TOF MS、MALDI-FT MS和LC MS/MS数据集进行程序检测。该方案允许将通过不同仪器获得的大量数据转换为若干商业和公开搜索引擎的格式。然后将文件提交给具有用户指定的搜索设置的搜索引擎进行蛋白质鉴定。我们的系统生物学研究所推出的mzXML格式的实施提供了一个共同的数据格式的好处,在不同的MS平台上获得的结果汇总,MS方法的比较分析和数据归档。我们增加了解释蛋白质自上而下串联质谱(BUPID-自上而下)和将数据库分配与蛋白质功能联系起来(STRAP)的功能。这些结果在ASMS和其他科学会议上以海报的形式展示; STRAP手稿于2010年初出版。目前正在开发另一个版本(STRAP-PTM),以分配和映射PTM。搜索算法Boston University Protein Identifier(BUPID)为使用MS数据进行蛋白质鉴定提供了一个鲁棒且准确的统计模型。该算法提供了一些重要的功能:1.使用对数似然比作为评分函数,该算法可以最好地区分正确分配的肽与不正确的分配。2.与传统的质量窗口相比,使用背景相关阈值匹配峰提供了更高的灵活性和准确性。3.与传统的数据库搜索引擎相比,统计模型提供了类似或更好的结果。我们使用对数似然比来计算样本中存在蛋白质的概率。该模型区分了两个假设:(1)H 0:光谱中的一组峰是由随机背景产生的;(2)HA:同一组峰是由对应于特定蛋白质的肽产生的。如果由蛋白质产生的峰的概率比由随机背景产生的峰的概率更显著,则将该峰包括在该组中。使用蛋白质的序列信息,通过其概率得分的E值对最终结果进行排名。我们比较了BUPID服务器和其他几个公共的基于Web的数据库搜索引擎的性能。最近的努力集中在解释自上而下的串联MS数据从LTQ轨道阱MS和ICR-FTMS仪器。最近在《国际质谱杂志》上发表了一份手稿,其中包括BUPID自上而下的使用。还发表了一篇介绍STRAP的论文,目前正在修订关于STRAP-PTM的手稿。
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. Primary support for the subproject and the subproject's principal investigator may have been provided by other sources, including other NIH sources. The Total Cost listed for the subproject likely represents the estimated amount of Center infrastructure utilized by the subproject, not direct funding provided by the NCRR grant to the subproject or subproject staff. With a multitude of different MS instrumentation and data analysis software platforms available for MS and proteomics, it becomes difficult to manipulate and manage various data sets. We are creating new software to process protein and peptide MS and MS/MS data, locate and assign post-translational modifications (PTMs) and to help place the results into biological context. The BUPID program was developed in C under Linux and made accessible to the main program through a CGI based web interface. The shell data conversion program was written to implement a user friendly GUI interface which may be operated in an unattended batch processing mode. Testing of the program was performed on existing MALDI-TOF MS, MALDI-FT MS and LC MS/MS data sets obtained in house. The program allowed the conversion of large volumes of data obtained on different instruments to the formats of several commercially and publicly available search engines. Files were then submitted for protein identification to the search engines with the search settings specified by the user. Our implementation of the mzXML format introduced by the Institute for Systems Biology afforded the benefits of a common data format for summation of results obtained on different MS platforms, comparative analysis of MS methodology and archiving of data. We added capabilities for interpretation of top-down tandem mass spectra of proteins (BUPID-top down) and for linking of database assignments to functionality of proteins (STRAP). These results were presented as posters at ASMS and other scientific meetings; the STRAP manuscript was published in early 2010. A further dedition (STRAP-PTM) is now being developed to assign and map PTMs. The search algorithm Boston University Protein Identifier (BUPID) provides a robust and accurate statistical model for protein identification using MS data. The algorithm offers a number of important features: 1. Using log-likelihood ratio as scoring function, the algorithm can best distinguish correctly assigned peptides from incorrect assignments. 2. Matching peaks with a background-dependent threshold offers more flexibility and accuracy than the traditional mass window. 3. The statistical model provides similar or better results with comparison to conventional database search engines. We use log-likelihood ratio to calculate the probability that a protein is present in the sample. The model distinguishes two hypotheses (1) H0: That a set of peaks in the spectrum is generated by the random background; and (2) HA: That the same set of peaks is generated by peptides corresponding to a specific protein. A peak is included in the set if the probability that it is produced by the protein is more significant than that it is otherwise produced by the random background. Final results are ranked by the E-value of their probability score using the sequence information of the protein. We have compared the performance of the BUPID server and several other public web-based database search engines. Recent efforts have focused on interpretation of top-down tandem MS data from the LTQ-Orbitrap MS and ICR-FTMS instruments. A manuscript that includes the use of BUPID Top-down was recently published in the Int. J. Mass Spectrom. A paper describing STRAP was also published and a manuscript on STRAP-PTM is now under revision.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Legacy Support During Closure of the Mass Spectrometry Resource for Biology and Medicine
  • 批准号:
    10204050
  • 项目类别:
  • 资助金额:
    $53.99万
  • 财政年份:
    2019
  • 负责人:
    Catherine E. Costello
  • 依托单位:
Legacy Support During Closure of the Mass Spectrometry Resource for Biology and Medicine
  • 批准号:
    9976561
  • 项目类别:
  • 资助金额:
    $70.81万
  • 财政年份:
    2019
  • 负责人:
    Catherine E. Costello
  • 依托单位:
Legacy Support During Closure of the Mass Spectrometry Resource for Biology and Medicine
  • 批准号:
    9810729
  • 项目类别:
  • 资助金额:
    $82.73万
  • 财政年份:
    2019
  • 负责人:
    Catherine E. Costello
  • 依托单位:
MALDI-TOF/TOF MS TO SUPPORT BIOMEDICAL RESEARCH
  • 批准号:
    8247392
  • 项目类别:
  • 资助金额:
    $59.0万
  • 财政年份:
    2012
  • 负责人:
    Catherine E. Costello
  • 依托单位:
海外基金