A graphical method for the analysis of statistical distributions into two normal components
A graphical method for the analysis of statistical distributions into two normal components
复制标题
用于分析两个正态分量的统计分布的图形方法
DOI:
10.1093/biomet/40.3-4.460
复制
发表时间:
1953
期刊:
影响因子:
2.7
通讯作者:
E. J. Preston
中科院分区:
文献类型:
--
作者:
E. J. Preston
Many frequency distributions occur in statistical practice which are probably due to the existence of two or more separate sub-universes within the universe under consideration, each of which is of approximately Normal form; for instance, the heights or weights of English men and women, or the intelligence quotients of professional and artisan engineers would probably have this property. If the sub-universes have different mean values, different variances and contain different numbers of individuals, a great variety of composite distributions will result. The existence of sub-universes is not usually obvious from the bimodal or multi-modal appearance of a frequency curve, since separate modes do not appear unless the separation of the means is considerable; it needs to be about three times the standard deviation of the components, if they are of comparable size. The following problem therefore arises: given a sample (preferably of several thousand individuals), selected from a universe suspected for some reason of having such components, can we determine the most probable nature of the sub-universes? Perhaps most important, is there a quick practical way of doing so? The'method of moments' was first used in a theoretical solution of the problem by K. Pearson (1894). Since then, the more efficient'method of maximum likelihood'has been developed by RA Fisher and others, and applied in particular to this problem by CR Rao (1948). Rao also gives, in the same paper, a rapid and elegant theoretical solution, by the method of moments, for the case of two components assumed to have equal variances. This depends upon the solution of a cubic equation, and appears to yield quite accurate results.However, it seems to us that all these contributions suffer from over-complexity and the necessity for lengthy calculations. The methods do indeed give the most accurate results possible from the available sample, but this degree of accuracy seems rather unnecessary in view of the unavoidable inherent errors due to sampling fluctuations and to the assumptions that the true components are normal and have equal variances. Thus there seems to be room for a more rapid, and much less laborious, graphical method.