Reliable identification of large numbers of candidate SNPs from public EST data

Reliable identification of large numbers of candidate SNPs from public EST data
复制标题

DOI:
10.1038/6851
复制
发表时间:
1999-03-01
期刊:
影响因子:
30.8
通讯作者:
Cassidy, AB
Cassidy, AB
中科院分区:
生物学1区
文献类型:
--
作者:
Buetow, KH;Edmonson, MN;Cassidy, AB

文献摘要

被引文献

相似文献

对人类基因组的高分辨率遗传分析有望提供对常见疾病易感性的深入了解。为了进行这种分析,将需要收集高通量、高密度的分析试剂。我们已经开发了一个多态性检测系统,使用公共域序列数据。这种检测系统被称为单核苷酸多态性流水线(SNP流水线)。SNPpipeline的分析核心由PHRED、PHRAP和DEMIGLACE三个组件组成。PHRED和PHRAP是序列分析套件的组成部分,开发用于执行大规模基因组所需的半自动分析(1,2)(由P.绿色提供)。使用这些信息学工具,检查冗余的原始表达序列标签(EST)数据,我们已经确定了3,000多个候选单核苷酸多态性(SNP)。一组192名候选人的经验验证研究表明,82%的识别10个中心的多态性人类(CEPH)个人的样本中的变化。我们的研究结果表明,现有的序列资源可以作为一个有价值的来源,为确定遗传变异。
High-resolution genetic analysis of the human genome promises to provide insight into common disease susceptibility. To perform such analysis will require a collection of high-throughput, high-density analysis reagents. We have developed a polymorphism detection system that uses public-domain sequence data. This detection system is called the single nucleotide polyrmorphism pipeline (SNPpipeline). The analytic core of the SNPpipeline is composed of three components: PHRED, PHRAP and DEMIGLACE. PHRED and PHRAP are components of a sequence analysis suite developed to perform the semi-automated analysis required for large-scale genomes(1,2) (provided courtesy of P. Green). Using these informatics tools, which examine redundant raw expressed sequence tag (EST) data, we have identified more than 3,000 candidate single-nucleotide polyrmorphisms (SNPs). Empiric validation studies of a set of 192 candidates indicate that 82% identify variation in a sample of ten Centre d'Etudes Polymorphism Humain (CEPH) individuals. Our results suggest that existing sequence resources may serve as a valuable source for identifying genetic variation.