Statistical analysis of molecular genetic data.

Statistical analysis of molecular genetic data.
复制标题

分子遗传数据的统计分析。

DOI:
10.1093/imammb/2.1.1
复制
发表时间:
1985
期刊:
IMA journal of mathematics applied in medicine and biology
影响因子:
--
通讯作者:
Weir,BS
Weir,BS
中科院分区:
--
文献类型:
--
作者:
Weir,BS

文献摘要

被引文献

相似文献

分子生物学的最新进展使遗传学家能够在DNA序列水平上收集数据,从而直接对基因进行观察,而不是对受基因影响的特征进行观察。如本文所示,遗传数据的这种新性质给统计分析带来了新的问题。在解释了重组DNA研究的方法学之后,分别考虑了限制性内切酶图谱和全序列数据的统计分析。限制性内切酶研究的初步观察是用酶消化DNA区域所产生的片段的数量和大小,因此有必要推断断裂点的位置。然后有必要进行基于模型的分析,以从这些断裂点的频率推断DNA序列的特征。这种限制图的一个重要应用是检测人类疾病基因,迄今为止最显著的成功是亨廷顿舞蹈病基因的染色体分配。完整的序列数据存在规模上的冲突问题。每个个体收集的信息量如此之大,以至于分析必须始终在计算机上进行,但在一个种群中研究的个体数量太少,无法评估种群之间和种群内的差异大小。统计问题还包括序列模式的检测和测试以及序列的比较。任何分析都必须结合观察到的相邻序列成员之间的高度关联。虽然国家数据库现在包含大约300万个序列元素,但人们对这些序列的生成或记录的错误率知之甚少。
Recent advances in molecular biology have made it possible for geneticists to collect data at the DNA sequence level, and so make observations on genes directly, instead of on characters affected by genes. This new nature of genetic data is posing new problems of statistical analyses, as is shown in this paper. After an explanation of the methodology of recombinant DNA studies, separate consideration is given to the statistical analyses of restriction-map and complete-sequence data.The initial observations from restriction-enzyme studies are of the number and sizes of the fragments produced by digesting a DNA region with the enzymes, and it is necessary to infer the location of the breakage points. Model-based analyses are then necessary to infer characteristics of the DNA sequence from the frequency of these break points. An important application of such restriction maps is the detection of human disease genes, with the most notable success to date being die chromosomal allocation of the gene for Huntington's chorea.Complete sequence data present conflicting problems of scale. The amount of information collected per individual is so large that analyses must always be done on a computer, but the number of individuals studied in a population is too low to allow the magnitudes of between- and within-population variation to be assessed. Statistical problems also include the detection and testing of sequence patterns and the comparison of sequences. Any analyses are going to have to incorporate the observed high levels of association between adjacent sequence members. While national data bases now contain about three million sequence elements, very little is known about the error rate in the generation or recording of these sequences.