Optimal genotype determination in highly multiplexed SNP data

Optimal genotype determination in highly multiplexed SNP data
复制标题

DOI:
10.1038/sj.ejhg.5201528
复制
发表时间:
2006-02-01
影响因子:
5.2
通讯作者:
Faham, M
Faham, M
中科院分区:
生物学2区
文献类型:
--
作者:
Moorhead, M;Hardenbol, P;Faham, M

文献摘要

被引文献

相似文献

高通量基因分型技术使大型关联研究成为可能。从原始信号强度开始的基因型测定工具需要自动化、健壮和灵活,以提供给定研究特定要求的最佳基因型测定。描述自定义基因分型研究性能的关键指标是分析转换、呼叫率和基因型准确性。这三个指标可以相互权衡。以高度复用的分子反转探针技术为例,我们描述了一种识别最佳权衡的方法。该方法包括:一种鲁棒的聚类算法和对大量数据过滤集的评估。聚类算法允许自动确定基因型。然后将许多不同的过滤器集应用于聚集的数据,并计算每个过滤器集产生的性能指标。这些性能指标与研究的能力有关,并为为特定研究选择最合适的过滤器集提供了一个框架。
High-throughput genotyping technologies that enable large association studies are already available. Tools for genotype determination starting from raw signal intensities need to be automated, robust, and flexible to provide optimal genotype determination given the specific requirements of a study. The key metrics describing the performance of a custom genotyping study are assay conversion, call rate, and genotype accuracy. These three metrics can be traded off against each other. Using the highly multiplexed Molecular Inversion Probe technology as an example, we describe a methodology for identifying the optimal trade-off. The methodology comprises: a robust clustering algorithm and assessment of a large number of data filter sets. The clustering algorithm allows for automatic genotype determination. Many different sets of filters are then applied to the clustered data, and performance metrics resulting from each filter set are calculated. These performance metrics relate to the power of a study and provide a framework to choose the most suitable filter set to the particular study.