Identification and analysis of small proteins and short open reading frame encoded peptides in Hep3B cell

Identification and analysis of small proteins and short open reading frame encoded peptides in Hep3B cell
复制标题

Hep3B 细胞中小蛋白和短开放阅读框编码肽的鉴定和分析

DOI:
10.1016/j.jprot.2020.103965
复制
发表时间:
2020
影响因子:
3.3
通讯作者:
Cuihong Wan
Cuihong Wan
中科院分区:
生物学2区
文献类型:
--
作者:
Bing Wang;Junhui Hao;Ni Pan;Zhiwei Wang;Yinxuan Chen;Cuihong Wan

文献摘要

相似文献

小分子蛋白质和短的开放阅读框编码肽(SEP)在生物学过程中起着重要的作用。然而,它们的注释或鉴定是具有挑战性的,部分原因是传统基因组注释管道的限制以及它们低丰度和低分子量的固有特性。为了发现和表征Hep 3B细胞系中的SEPs,我们通过结合不同的肽提取和分离方法开发了优化的肽组分析法。肽组学中的有机溶剂沉淀法显示出对低分子蛋白质或肽的富集的促进作用,并且数据清楚地显示出降低样品复杂性的有益效果,从而产生高质量的MS/MS谱。此外,不同的策略表现出良好的互补性,提高小蛋白的总量和它们的序列覆盖率。总共鉴定了1192个少于100个氨基酸的蛋白质,包括271个新发现的SEP,这些SEP在OpenProt数据库中被注释,其中147个SEP由ncRNA或lincRNA编码。这项工作的结果提供了强有力的证据,迄今为止,人类蛋白质组比以前认识到的更复杂,这将是一个有益的发现蛋白质没有功能annotation.SignificanceIn这项工作中,方法进行了优化,以确定SEPs在Hep 3B。有机溶剂沉淀促进了低分子蛋白质或肽的富集,数据清楚地显示了降低样品复杂性的有益效果,从而获得高质量的MS/MS谱。不同策略在提高小蛋白总量及其序列覆盖率方面表现出良好的互补性。总共鉴定了1192个少于100个氨基酸的蛋白质,包括271个新发现的SEP,这些SEP在OpenProt数据库中被注释,其中147个SEP由ncRNA或lincRNA编码。此外,从uORF中获得的22个SEP可能具有潜在的翻译调控功能,149个新鉴定的SEP具有已知的功能结构域或跨物种保守性。这项工作的结果为人类基因组中被忽视的区域的编码潜力提供了强有力的证据,并可能为肿瘤生物学提供更多的见解。
The small proteins and short open reading frames encoded peptides (SEPs) are of fundamental importance because of their essential roles in biological processes. However, the annotation or identification of them is challenging, in part owing to the limitation of the traditional genome annotation pipeline and their inherent characteristics of low abundance and low molecular weight. To discover and characterize SEPs in Hep3B cell line, we developed an optimized peptidomic assay by combining different peptide extraction and separation methods. The organic solvent precipitation method in peptidomic showed promotion in the enrichment of low molecular proteins or peptides, and the data clearly showed a beneficial effect from the reduction of sample complexity, resulting in high-quality MS/MS spectra. Furthermore, different strategies exhibited good complementarity in improving the total amount of small proteins and their sequence coverage. In total, 1192 proteins within less than 100 amino acids were identified, including 271 newly discovered SEPs that been annotated in the OpenProt database and 147 SEPs of them encoded from ncRNA or lincRNA. Results in this work provide robust evidence to date that the human proteome is more complicated than previously appreciated, and this will be a benefit to discoveries of proteins without function annotation.SignificanceIn this work, methods were optimized to identify SEPs in Hep3B. The organic solvent precipitation presents promotion in enrichment of low molecular proteins or peptides, and the data clearly showed a beneficial effect from the reduction of sample complexity, resulting in high quality MS/MS spectra. Different strategies exhibited good complementarity in improving total amount of small proteins and their sequence coverage. In total, 1192 proteins within less than 100 amino acids were identified, including 271 newly discovered SEPs that been annotated in the OpenProt database and 147 SEPs of them encoded from ncRNA or lincRNA. Furthermore, 22 SEPs generated from the uORF may has potential effect in translation control, and 149 newly identified SEPs have known functional domains or cross-species conservation. Results in this work present robust evidence for the coding potential of the ignored region of human genomes and may provide additional insights into tumor biology.