Toward high-throughput genotyping: Dynamic and automatic software for manipulating large-scale genotype data using fluorescently labeled dinucleotide markers

Toward high-throughput genotyping: Dynamic and automatic software for manipulating large-scale genotype data using fluorescently labeled dinucleotide markers
复制标题

DOI:
10.1101/gr.159701
复制
发表时间:
2001-07-01
期刊:
影响因子:
7
通讯作者:
Deng, HW
Deng, HW
中科院分区:
生物学1区
文献类型:
--
作者:
Li, JL;Deng, HY;Deng, HW

文献摘要

被引文献

相似文献

为了有效地操作大量的基因型数据产生的荧光标记的二核苷酸标记,我们开发了一个Microsoft Access数据库管理系统,命名为GenoDB。GenoDB提供了几个优势。首先,它适应了基因分型过程中基因型数据积累的动态性质;一些数据需要通过重复实验室程序进行确认或替换。通过使用GenoDB,原始基因型数据可以轻松连续地导入,并在大型项目中可能持续很长一段时间的基因分型过程中纳入数据库。其次,几乎所有的程序都是自动的,包括不同技术人员从同一凝胶读取的原始数据的自动比较,来自交叉运行或跨平台的等位基因片段大小数据之间的自动调整,等位基因的自动分组,以及基因型数据的自动编译,以用于合适的程序来执行家系中的遗传检查。第三,GenoDB提供了跟踪电泳凝胶文件的功能,以定位任何所得基因型数据的凝胶或样品来源,这对于复查原始数据和最终数据的一致性以及指导重复实验非常有帮助。此外,GenoDB用户友好的图形界面使处理大量数据的劳动密集程度大大降低。此外,GenoDB具有内置的机制来检测一些基因分型错误,并评估基因型数据的质量,然后汇总在GenoDB自动生成的统计报告中。GenoDB可以轻松处理超过500,000个基因型数据条目,这对于典型的全基因组连锁研究来说是绰绰有余的。如果需要同时处理更大量基因型数据的能力,我们为GenoDB开发的模块和程序可以扩展到其他数据库平台,例如Microsoft SQL server。
To efficiently manipulate large amounts of genotype data generated with fluorescently labeled dinucleotide markers, we developed a Microsoft Access database management system, named GenoDB. GenoDB offers several advantages. First, it accommodates the dynamic nature of the accumulations of genotype data during the genotyping process; some data need to be confirmed or replaced by repeat lab procedures. By using GenoDB, the raw genotype data can be imported easily and continuously and incorporated into the database during the genotyping process that may continue over an extended period of time in large projects. Second, almost all of the procedures are automatic, including autocomparison of the raw data read by different technicians from the same gel, autoadjustment among the allele fragment-size data from cross-runs or cross-platforms, autobinning of alleles, and autocompilation of genotype data for suitable programs to perform inheritance check in pedigrees. Third, GenoDB provides functions to track electrophoresis gel files to locate gel or sample sources for any resultant genotype data, which is extremely helpful for double-checking consistency of raw and final data and for directing repeat experiments. In addition, the user-friendly graphic interface of GenoDB renders processing of large amounts of data much less labor-intensive. Furthermore, GenoDB has built-in mechanisms to detect some genotyping errors and to assess the quality of genotype data that then are summarized in the statistic reports automatically generated by GenoDB. The GenoDB can easily handle >500,000 genotype data entries, a number more than sufficient for typical whole-genome linkage studies. The modules and programs we developed for the GenoDB can be extended to other database platforms, such as Microsoft SQL server, if the capability to handle still greater quantities of genotype data simultaneously is desired.