Different evolutionary patterns of SNPs between domains and unassigned regions in human protein-coding sequences.

Different evolutionary patterns of SNPs between domains and unassigned regions in human protein-coding sequences.
复制标题

人类蛋白质编码序列中结构域和未指定区域之间 SNP 的不同进化模式

DOI:
10.1007/s00438-016-1170-7
复制
发表时间:
2016-06
期刊:
Molecular genetics and genomics : MGG
影响因子:
--
通讯作者:
Lin K
Lin K
中科院分区:
其他
文献类型:
--
作者:
Pang E;Wu X;Lin K

文献摘要

相似文献

蛋白质进化在每个基因组的进化中起着重要的作用。由于它们的功能性质,一般来说,它们的大多数部分或位点都受到不同的选择性限制,特别是通过纯化选择。大多数以前的蛋白质进化研究认为个别蛋白质的整体或比较蛋白质编码序列与非编码序列。较少关注给定基因组的每个蛋白质内不同部分的进化。为此,基于所有人类蛋白质的PfamA注释,每个蛋白质序列可以分成两部分:结构域或未分配区域。利用这一原理,根据两种分类绘制了1000基因组计划蛋白质编码序列中的单核苷酸多态性(SNP):蛋白质结构域内发生的SNP和未分配区域内发生的SNP。通过这些分类,我们发现:在域内的同义SNPs的密度显着大于未分配区域内的同义SNPs的密度,然而,非同义SNPs的密度显示相反的模式。我们还发现在结构域和未分配区域都存在纯化选择的特征。此外,对结构域的选择性强度显著大于未分配区域。此外,在所有的人类蛋白质序列中,有117个PfamA结构域没有发现SNPs,我们的结果突出了蛋白质结构域的一个重要方面,可能有助于我们理解蛋白质的进化。本文的在线版本(doi:10.1007/s 00438 -016-1170-7)包含补充材料,可供授权用户使用。
Protein evolution plays an important role in the evolution of each genome. Because of their functional nature, in general, most of their parts or sites are differently constrained selectively, particularly by purifying selection. Most previous studies on protein evolution considered individual proteins in their entirety or compared protein-coding sequences with non-coding sequences. Less attention has been paid to the evolution of different parts within each protein of a given genome. To this end, based on PfamA annotation of all human proteins, each protein sequence can be split into two parts: domains or unassigned regions. Using this rationale, single nucleotide polymorphisms (SNPs) in protein-coding sequences from the 1000 Genomes Project were mapped according to two classifications: SNPs occurring within protein domains and those within unassigned regions. With these classifications, we found: the density of synonymous SNPs within domains is significantly greater than that of synonymous SNPs within unassigned regions; however, the density of non-synonymous SNPs shows the opposite pattern. We also found there are signatures of purifying selection on both the domain and unassigned regions. Furthermore, the selective strength on domains is significantly greater than that on unassigned regions. In addition, among all of the human protein sequences, there are 117 PfamA domains in which no SNPs are found. Our results highlight an important aspect of protein domains and may contribute to our understanding of protein evolution. The online version of this article (doi:10.1007/s00438-016-1170-7) contains supplementary material, which is available to authorized users.