A genotype calling algorithm for the Illumina BeadArray platform.

A genotype calling algorithm for the Illumina BeadArray platform.
复制标题

DOI:
10.1093/bioinformatics/btm443
复制
发表时间:
2007-10-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Clark TG
Clark TG
中科院分区:
其他
文献类型:
--
作者:
Teo YY;Inouye M;Small KS;Gwilliam R;Deloukas P;Kwiatkowski DP;Clark TG

文献摘要

参考文献

被引文献

相似文献

大规模基因分型依赖于使用无监督的自动调用算法来分配杂交数据的基因型。最近已经建立了许多这样的调用算法用于Affyssin基因芯片基因分型技术。在这里,我们提出了一个快速和准确的基因型调用算法的Illumina BeadArray基因分型平台。随着该技术向同时测定数百万种遗传多态性的方向发展,需要一种集成且易于使用的软件来调用基因型。我们引入了一种基于模型的基因型识别算法,该算法不依赖于先前的训练数据,也不需要计算密集型程序。该算法可以同时将基因型分配给来自数千个个体的杂交数据,并将多个个体的信息集中起来以提高识别率。该方法可以通过识别最佳坐标来初始化算法来适应杂交强度的变化,所述杂交强度的变化导致基因型云的位置的显著偏移。通过结合扰动分析的过程,我们可以获得测量分配的基因型调用的稳定性的质量度量。我们表明,这种质量指标可用于识别具有低呼叫率和准确性的SNP。这里描述的算法的C++可执行文件可通过作者的请求获得。
Large-scale genotyping relies on the use of unsupervised automated calling algorithms to assign genotypes to hybridization data. A number of such calling algorithms have been recently established for the Affymetrix GeneChip genotyping technology. Here, we present a fast and accurate genotype calling algorithm for the Illumina BeadArray genotyping platforms. As the technology moves towards assaying millions of genetic polymorphisms simultaneously, there is a need for an integrated and easy-to-use software for calling genotypes. We have introduced a model-based genotype calling algorithm which does not rely on having prior training data or require computationally intensive procedures. The algorithm can assign genotypes to hybridization data from thousands of individuals simultaneously and pools information across multiple individuals to improve the calling. The method can accommodate variations in hybridization intensities which result in dramatic shifts of the position of the genotype clouds by identifying the optimal coordinates to initialize the algorithm. By incorporating the process of perturbation analysis, we can obtain a quality metric measuring the stability of the assigned genotype calls. We show that this quality metric can be used to identify SNPs with low call rates and accuracy. The C++ executable for the algorithm described here is available by request from the authors.
DOI: 10.1093/bioinformatics/btm131
发表时间: 2007-06-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Xiao, Yuanyuan;Segal, Mark R.;Yeh, Ru-Fang
通讯作者: Yeh, Ru-Fang
DOI: 10.1093/biostatistics/kxl042
发表时间: 2007-04-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Carvalho, Benilton;Bengtsson, Henrik;Irizarry, Rafael A.
通讯作者: Irizarry, Rafael A.
DOI: 10.1038/sj.ejhg.5201528
发表时间: 2006-02-01
影响因子: 5.2
作者:
Moorhead, M;Hardenbol, P;Faham, M
通讯作者: Faham, M
DOI: 10.1038/ng2032
发表时间: 2007-05-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Rioux, John D.;Xavier, Ramnik J.;Brant, Steven R.
通讯作者: Brant, Steven R.
DOI: 10.2217/14622416.7.4.641
发表时间: 2006-06-01
期刊: PHARMACOGENOMICS
影响因子: 2.1
作者:
Gunderson, KL;Kuhn, KM;Shen, R
通讯作者: Shen, R