Proteomics Standards Initiative Extended FASTA Format

Proteomics Standards Initiative Extended FASTA Format
复制标题

DOI:
10.1021/acs.jproteome.9b00064
复制
发表时间:
2019-06-01
影响因子:
4.4
通讯作者:
Deutsch, Eric W.
Deutsch, Eric W.
中科院分区:
生物学2区
文献类型:
--
作者:
Binz, Pierre-Alain;Shofstahl, Jim;Deutsch, Eric W.

文献摘要

被引文献

相似文献

基于质谱的蛋白质组学能够对蛋白质进行高通量鉴定和定量,包括生物样品中的序列变异和翻译后修饰 (PTM)。然而,大多数工作流程要求将此类变化包含在用于分析数据的搜索空间中,而对于大多数分析工具来说,这样做仍然具有挑战性。为了促进已知序列变异和 PTM 的搜索,蛋白质组学标准倡议 (PSI) 设计并实施了 PSI 扩展 FASTA 格式 (PEFF)。 PEFF 基于非常流行的 FASTA 格式,但添加了一种统一机制,用于编码更多有关序列集合以及单个条目的元数据,包括支持编码已知序列变体、PTM 和蛋白质形式。该格式几乎向后兼容,因此,现有的 FASTA 解析器只需很少的更改或无需更改即可将 PEFF 文件读取为 FASTA 文件,尽管不支持 PEFF 的任何额外功能。 PEFF 由完整的规范文档、受控词汇术语、一组示例文件、软件库和文件验证器定义。流行的软件和资源开始支持 PEFF,包括序列搜索引擎 Comet 以及知识库 neXtProt 和 UniProtKB。 PEFF 的广泛实施预计将通过提供编码蛋白质序列及其已知变异的标准化机制,进一步实现蛋白质基因组学和自上而下的蛋白质组学应用。所有相关文档,包括详细的文件格式规范和示例文件,均可从 http://www.psidev.info/peff 获取。
Mass-spectrometry-based proteomics enables the high-throughput identification and quantification of proteins, including sequence variants and post-translational modifications (PTMs) in biological samples. However, most workflows require that such variations be included in the search space used to analyze the data, and doing so remains challenging with most analysis tools. In order to facilitate the search for known sequence variants and PTMs, the Proteomics Standards Initiative (PSI) has designed and implemented the PSI extended FASTA format (PEFF). PEFF is based on the very popular FASTA format but adds a uniform mechanism for encoding substantially more metadata about the sequence collection as well as individual entries, including support for encoding known sequence variants, PTMs, and proteoforms. The format is very nearly backward compatible, and as such, existing FASTA parsers will require little or no changes to be able to read PEFF files as FASTA files, although without supporting any of the extra capabilities of PEFF. PEFF is defined by a full specification document, controlled vocabulary terms, a set of example files, software libraries, and a file validator. Popular software and resources are starting to support PEFF, including the sequence search engine Comet and the knowledge bases neXtProt and UniProtKB. Widespread implementation of PEFF is expected to further enable proteogenomics and top-down proteomics applications by providing a standardized mechanism for encoding protein sequences and their known variations. All the related documentation, including the detailed file format specification and example files, are available at http://www.psidev.info/peff.