Human copy number polymorphic genes

Human copy number polymorphic genes
复制标题

DOI:
10.1159/000184713
复制
发表时间:
2008-01-01
影响因子:
1.7
通讯作者:
Eichler, E. E.
Eichler, E. E.
中科院分区:
生物学4区
文献类型:
--
作者:
Bailey, J. A.;Kidd, J. M.;Eichler, E. E.

文献摘要

被引文献

相似文献

最近在人类群体中的大规模基因组研究已经确定了大量的基因组区域为拷贝数变异(CNV)。由于这些CNV区域经常与基因组的编码区重叠,已经产生了大量潜在拷贝数多态基因的列表,这些基因是疾病关联的候选基因。然而,目前大多数关于正常基因变异的数据是使用BAC或SNP微阵列产生的,这些微阵列缺乏精度,特别是关于外显子。为了解决这个问题,我们通过设计一个专门针对外显子的定制寡核苷酸微阵列,在9个特征良好的HapMap个体中评估了从现有研究中定义的2790个候选CNV基因。利用外显子阵列比较基因组杂交(ACGH),我们检测到255个(9%)的候选CNV是真实的CNV,其中134个有全基因变异的证据。个体与对照个体在拷贝数上平均相差100个基因座。部分和全基因CNV都与节段性重复(分别为55%和71%)以及阳性选择区域密切相关。我们使用融合末端序列对(Fosmidend Sequence Pair,ESP)结构变异图对这些相同的个体确认了37%的全基因CNV。如果我们修改末端序列对作图策略,包括低序列同源性ESP(98-99.5%)和具有外翻方向的ESP,我们可以捕获82%的缺失基因,从而更完整地确定重复基因中的结构变异。我们的结果表明,片段复制是大多数全长拷贝数多态基因的来源,大多数变异基因以串联复制的形式组织,其中很大一部分基因将代表序列多样性水平超过等位基因变异阈值的平行基因。此外,这些数据提供了一组靶向CNV基因,丰富了可能与由于拷贝数变化而导致的人类表型差异相关的区域,并为未来的关联研究提供了拷贝数响应性寡核苷酸探针的来源。版权所有(C)2009年S.Karger AG,巴塞尔
Recent large-scale genomic studies within human populations have identified numerous genomic regions as copy number variant (CNV). As these CNV regions often overlap coding regions of the genome, large lists of potentially copy number polymorphic genes have been produced that are candidates for disease association. Most of the current data regarding normal genic variation, however, has been generated using BAC or SNP microarrays, which lack precision especially with respect to exons. To address this, we assessed 2,790 candidate CNV genes defined from available studies in nine well-characterized HapMap individuals by designing a customized oligonucleotide microarray targeted specifically to exons. Using exon array comparative genomic hybridization (aCGH), we detected 255 (9%) of the candidates as true CNVs including 134 with evidence of variation over the entire gene. Individuals differed in copy number from the control by an average of 100 gene loci. Both partial- and whole-gene CNVs were strongly associated with segmental duplications (55 and 71%, respectively) as well as regions of positive selection. We confirmed 37% of the whole-gene CNVs using the fosmid end sequence pair (ESP) structural variation map for these same individuals. If we modify the end sequence pair mapping strategy to include low-sequence identity ESPs (98-99.5%) and ESPs with an everted orientation, we can capture 82% of the missed genes leading to more complete ascertainment of structural variation within duplicated genes. Our results indicate that segmental duplications are the source of the majority of full-length copy number polymorphic genes, most of the variant genes are organized as tandem duplications, and a significant fraction of these genes will represent paralogs with levels of sequence diversity beyond thresholds of allelic variation. In addition, these data provide a targeted set of CNV genes enriched for regions likely to be associated with human phenotypic differences due to copy number changes and present a source of copy number responsive oligonucleotide probes for future association studies. Copyright (c) 2009 S. Karger AG, Basel