A DNA polymorphism discovery resource for research on human genetic variation

A DNA polymorphism discovery resource for research on human genetic variation
复制标题

DOI:
10.1101/gr.8.12.1229
复制
发表时间:
1998-12-01
期刊:
影响因子:
7
通讯作者:
Chakravarti, A
Chakravarti, A
中科院分区:
生物学1区
文献类型:
--
作者:
Collins, FS;Brooks, LD;Chakravarti, A

文献摘要

被引文献

相似文献

随着在全基因组范围内发现DNA序列变异的方法的改进,鉴定赋予人类常见疾病易感性或抗性的基因将变得越来越可行(柯林斯等人,1997; Landegren等人,1998; Wang等人,1998)。为了促进DNA序列变异的发现,美国国立卫生研究院的国家人类基因组研究所(NHGRI)与疾病控制和预防中心,国家环境健康科学研究所和几个独立的研究人员合作,已经从450名来自世界所有主要地区的美国居民的样本中收集了DNA多态性发现资源。这一DNA多态性发现资源对于发现人类遗传变异将具有巨大的价值,其他后续研究可能与健康和疾病有关。迄今为止,在发现导致疾病风险的基因方面取得的大多数成功都是针对由单基因引起的高度渗透性疾病,如囊性纤维化(Kerem et al. 1989; Rommens et al. 1989)。为了定位影响这些罕见疾病的基因,研究人员对家族进行连锁分析,这需要跨越整个人类基因组的300-500个高度信息化的遗传标记。然而,要找到导致糖尿病、心脏病、癌症和精神疾病等常见疾病风险的基因要困难得多,因为这些表型受到多个基因的影响,每个基因的影响都很小;环境因素也很重要。与对家族进行连锁分析相比,对许多受影响和未受影响的个体进行关联分析可能更有效,这将需要分布在整个基因组中的数十万个变体(Risch和Merikangas 1996)。如此大量的变体目前不可用。DNA多态性发现资源旨在促进他们的发现。人类中大约90%的序列变异是DNA单碱基的差异,称为单核苷酸多态性(SNP)。基因编码区(cSNPs)或调控区的SNPs比其他地方的SNPs更可能导致功能差异。虽然大多数SNP不影响基因功能,但大量作图的SNP作为整个基因组的标记物对于发现确实影响基因功能的SNP将是有价值的,因为预期在人类基因组的许多区域中发现数十至数百个SNP的连锁不平衡。SNPs和cSNPs都可以通过DNA多态性发现资源进行鉴定。当比较两个随机染色体时,它们在0.1/1000个核苷酸处不同(Kwok et al. 1996)。当筛选40个个体的所有染色体时,在人类DNA的30亿个碱基中,预计将发现约1700万个SNP。预计这些SNP中只有一小部分位于编码区,因为编码区占基因组的约5%,且不太可能具有SNP(Nickerson et al. 1998)。因此,cSNP的数量估计为1500,000,平均每个基因约6个。
Identifying the genes conferring susceptibility or resistance to common human diseases should become increasingly feasible with improved methods for finding DNA sequence variants on a genome-wide scale (Collins et al. 1997; Landegren et al. 1998; Wang et al. 1998). To facilitate the discovery of DNA sequence variants, the National Human Genome Research Institute (NHGRI) of NIH, working with the Centers for Disease Control and Prevention, the National Institute of Environmental Health Sciences, and several individual investigators, has assembled a DNA Polymorphism Discovery Resource of samples from 450 US residents with ancestry from all the major regions of the world. This DNA Polymorphism Discovery Resource will be immensely valuable for the discovery of human genetic variation, which other follow-up studies can relate to health and disease. Most successes so far in finding genes that contribute to disease risk have been for highly penetrant diseases caused by single genes, such as cystic fibrosis (Kerem et al. 1989; Rommens et al. 1989). To locate genes affecting these rare disorders, researchers perform linkage analysis on families, which requires 300–500 highly informative genetic markers spanning the entire human genome. However, it has been considerably harder to locate the genes contributing to the risk of common diseases such as diabetes, heart disease, cancers, and psychiatric disorders, because these phenotypes are affected by multiple genes, each with small effect; environmental contributions are also important. Instead of linkage analysis on families it may be much more efficient to perform association analysis on many affected and unaffected individuals, which would require hundreds of thousands of variants spread over the entire genome (Risch and Merikangas 1996). Such a large number of variants is currently not available. The DNA Polymorphism Discovery Resource is designed to promote their discovery. About 90% of sequence variants in humans are differences in single bases of DNA, called single nucleotide polymorphisms (SNPs). SNPs in the coding regions of genes (cSNPs) or in regulatory regions are more likely to cause functional differences than SNPs elsewhere. Although most SNPs do not affect gene function, a large number of mapped SNPs will be valuable as markers throughout the genome for finding SNPs that do affect gene function, as linkage disequilibrium over tens to hundreds of kilobases is expected to be found in many regions of the human genome. Both SNPs and cSNPs can be identified by using the DNA Polymorphism Discovery Resource. When two random chromosomes are compared, they differ at∼ 1⁄ 1000 nucleotides (Kwok et al. 1996). When all chromosomes from 40 individuals are screened, about 17 million SNPs are expected to be found, out of the 3 billion bases in human DNA. Only a small proportion of these SNPs are expected to be in coding regions, as coding regions are∼ 5% of the genome and are less likely to have SNPs (Nickerson et al. 1998). Thus the number of cSNPs is estimated to be∼ 500,000, an average of about 6 per gene.