Surnames as neutral alleles: observations in Sardinia.

Surnames as neutral alleles: observations in Sardinia.
复制标题

DOI:
--
复制
发表时间:
1983-05
期刊:
影响因子:
--
通讯作者:
G. Zei;C. Guglielmino;E. Siri;A. Moroni;L. Cavalli-Sforza
G. Zei;C. Guglielmino;E. Siri;A. Moroni;L. Cavalli-Sforza
中科院分区:
生物学4区
文献类型:
--
作者:
G. Zei;C. Guglielmino;E. Siri;A. Moroni;L. Cavalli-Sforza

文献摘要

被引文献

相似文献

姓氏可以被认为是一个位点的等位基因,通常是父系遗传的。大量的姓氏使得它们对于评估人群之间的亲属关系非常有用,即使这种亲属关系的估计是有限的时间深度。来自撒丁岛的一组数据显示中性等位基因的KarlinMcGregor分布与Fisher的对数分布非常吻合。后者可以由前者衍生而来;如果N很大,它实际上是无法区分的。它更容易计算,唯一要估计的参数v可以通过一个简单的公式得到,v表示突变加迁移。在11个教区中,8个显示出非常合适的姓氏,3个显示姓氏在只代表一次的姓氏类别中过量。正如与发展统计数据的相关性所显示的那样,这很可能是由于最近在工业化地区纳入了独特的姓氏。姓氏亲属关系(同姓)的对数随教区之间的地理距离呈非线性下降,至少部分是由于不同频率姓氏的斜率的异质性。树分析结果与主成分分析结果吻合较好;它们显示出主要的南北分化,与岛屿的长轴相对应,并与遗传分化的数据一致。东西差异在某种程度上不那么重要,正如与第二个主成分相关的变异量较小所显示的那样,第三个主成分与工业发展程度相关。通过认识到姓氏通常具有父系遗传,但在其他方面表现得像单倍体遗传特征,姓氏的研究方法可以获得实质性的成果。至少对于大多数姓氏来说,在繁殖或死亡率方面不太可能有任何严重的偏见,因此它们可以被认为是“选择性中性的”。如果这是正确的,那么姓氏就可以用遗传学中中性等位基因的研究方法来分析。因此,在形式上,每个姓氏都可以被认为是一个基因位点的等位基因;“等位基因”的数量越多,基因座的信息就越多,因此,从原则上讲,姓氏可以为人们的起源、混合和迁移模式提供极其丰富的信息。当然,与真正的基因位点相比,它们的局限性在于它们只能告诉我们雄性迁徙的情况。人类起源信息的时间深度由标记的起源时间决定:意大利帕维亚的遗传生物化学和进化研究所;意大利帕尔马的生态研究所;加州斯坦福大学的遗传学系;《人类生物学》,1983年5月,vol . 55, No. 2, pp. 357-365。©韦恩州立大学出版社,1983此内容下载自157.55.39.253周三,2016年6月8日05:13:38 UTC所有使用受http://about.jstor.org/terms 358 G. Zei, С。Я。古列尔米诺(Guglielmino)、E. Siri、A.莫罗尼(A. Moroni)和L. L.卡瓦利-斯福尔扎(L. L. Cavalli-Sforza),就基因而言,通常是几千年或几万年,但就姓氏而言,很少早于中世纪晚期,至少在欧洲是这样。在本文中,我们分析了一组在撒丁岛未公布的近亲婚姻普查中获得的数据。所研究的姓氏属于近亲婚姻中的丈夫;由于近亲的同姓性通常很高,妻子被排除在外。近亲婚姻(N = 48,470),主要是在第一或第二表兄弟之间。1930-1959年,撒丁岛所有11个教区的主教都要求分配(图1)。从某种程度上说,近亲交配并不是人口的随机样本,这种选择可能会对我们的结果造成一些偏差。然而,在意大利的另一个地区进行的测试显示,在姓氏分布上没有明显的差异,因此我们的结果可以被认为是所有姓氏的合理代表。我们首先研究了中性等位基因的理论分布与在给定数量的个体中所代表的姓氏频率分布的拟合性。然后,我们将通过拟合这些理论分布获得的移民频率与从其他来源获得的频率进行了比较。其次,根据现有的遗传亲缘关系与地理距离的关系理论,我们研究了姓氏的亲缘关系估计作为地理距离的函数。最后,应用主成分分析和系统发育分析等群体遗传分析的标准方法对亲缘关系数据进行分析。中性等位基因的分布Karlin和McGregor(1967)给出了中性等位基因在N个(单倍体)个体的群体中的分布,每个个体携带一个不同的等位基因(在我们的例子中是姓氏),服从随机死亡的过程。死亡个体被携带相同等位基因(相同姓氏)的新个体所取代,除非突变以概率v的方式发生。Yasuda等人(1974)表明,这种分布与帕尔马山谷观察到的姓氏数据非常吻合,在帕尔马山谷可以获得移民数据。由于v代表姓氏突变(罕见)加上移民(频繁)的总和,我们将在接下来的内容中假设v估计移民。在本文中,我们还采用了Fisher等人(1943)首先给出的对数分布来计算昆虫样本中的物种丰度,而不是计算起来更麻烦的KarlinMcGregor分布。此内容于2016年6月8日星期三05:13:38 UTC从157.55.39.253下载,所有使用http://about.jstor.org/terms姓氏作为生物标记359图1。撒丁岛的教区很容易表明,Fisher分布可以从KarlinMcGregor分布导出为N»100,这两个分布已经非常吻合。使用费雪分布,在N个样本中,由个体代表的姓氏的数量预计为
Surnames can be considered as alleles of a locus which are usually transmitted patrilineally. The great abundance of surnames makes them very useful for evaluating kinship between populations, even if such kinship estimates are of limited time depth. A set of data from the island of Sardinia shows very good agreement with the KarlinMcGregor distribution of neutral alleles, and also with Fisher's logarithmic distribution. The latter can be derived from the former; it is practically indistinguishable from it if N is large. It is easier to compute and the only parameter to estimate, v, can be obtained by a simple formula, v measures mutation plus immigration. Of the 11 dioceses, eight showed a very good fit and three showed an excess of surnames in the class of surnames represented only once. This is most probably due to the recent incorporation of the unique surnames in industrialized areas, as the correlations with statistics of development show. The logarithms of surname kinship (isonymy) show a nonlinear decrease with geographic distance between dioceses, at least in part due to heterogeneity of the slopes for surnames of different frequency. Tree analysis and principal components of isonymy are in good agreement; they show a major north-south differentiation, corresponding to the major axis of the island, and in agreement with data on genetic differentiation. East-west divergence is somewhat less important, as shown by the smaller amount of variation associated with the second principal component, and the third principal component is correlated with degree of industrial development. Methods of study of surnames can gain substantially by the recognition that they usually have patrilineal transmission, but otherwise behave like a haploid genetic trait. At least for most surnames there is not likely to be any serious bias in reproduction or mortality, so that they can be considered as "selectively neutral. " If this is correct, surnames can be analyzed with the methods derived for the study of neutral alleles in genetics. Thus, formally, every surname can be considered as an allele of a genetic locus; the greater the number of "alleles," the greater the information of the locus, and therefore surnames can be, in principle, extremely informative on origins, admixture, and migration patterns of people. In comparison with true genetic loci, of course, they are limited by the fact that they inform us only on male migration. The time depth of information on origins of people is determined by the time of origin of the markers instituto di Genetica Biochimica ed Evoluzionistica, C.N.R., Pavia, Italy 2Instituto di Ecologia, Parma, Italy 3Department of Genetics, Stanford University, Stanford, CA 94305 Human Biology, May 1983, Voi. 55, No. 2, pp. 357-365. © Wayne State University Press, 1983 This content downloaded from 157.55.39.253 on Wed, 08 Jun 2016 05:13:38 UTC All use subject to http://about.jstor.org/terms 358 G. Zei, С. Я. Guglielmino, E. Siri, A. Moroni and L. L. Cavalli-Sforza themselves, which in the case of genes is usually of thousands and tens of thousands of years, but in the case of surnames is rarely earlier than the late Middle Ages, at least in Europeans. In the present paper we analyzed a set of data obtained in an unpublished census of consanguineous marriages in Sardinia. The surnames studied belong to the husbands in consanguineous marriages; because of the usually high isonymy of consanguineous mates, wives were excluded. Consanguineous marriages (N = 48,470), mostly between first or second cousins, were used. Dispensations were requested from bishops in all 11 dioceses of Sardinia (Figure 1) for the years 1930-1959. To the extent that consanguineous matings are not a random sample of the population, this selection may impose some bias on our results. However, tests in another part of Italy where both consanguineous and total matings were available showed no important discrepancies in surname distributions, so that our results can be considered as reasonably representative of all surnames. We first studied the fit of theoretical distributions of neutral alleles to the observed frequency distributions of surnames represented in a given number of individuals. We then compared the immigration frequencies obtained by fitting such theoretical distributions with those obtained from other sources. Next, in the light of current theories regarding the relationship between genetic kinship and geographic distance, we studied kinship estimates obtained from surnames as a function of geographic distance. Finally, we applied other standard methods of population genetic analysis, such as principal component and phylogenetic analysis, to the kinship data. Distribution of Neutral Alleles Karlin and McGregor (1967) have given the distribution of neutral alleles that are expected in a population of N (haploid) individuals each carrying one of к different alleles (surnames, in our case), subject to a process of death at random. Dead individuals are replaced by new ones carrying the same allele (the same surname) unless a mutation occurs with probability v. It was shown by Yasuda et al. (1974) that this distribution fits very well the observed surname data of the Parma Valley where it was possible to obtain immigration data. As v represents the sum of surname mutation (which is rare), plus immigration (which is frequent), we will assume in what follows that v estimates immigration. In the present paper we have also employed the logarithmic distribution first given by Fisher et al. (1943) for species abundance in insect samples instead of the KarlinMcGregor distribution which is more cumbersome to compute. One can This content downloaded from 157.55.39.253 on Wed, 08 Jun 2016 05:13:38 UTC All use subject to http://about.jstor.org/terms Surnames as Biological Markers 359 Fig. 1. Dioceses of the island of Sardinia easily show that the Fisher distribution can be derived from the KarlinMcGregor distribution as N » 100 the two distributions are already in very good agreement. Using Fisher s distribution, the number of surnames represented by к individuals in a sample of N is expected to be