The analysis of names from the 2011 Census of Population
The analysis of names from the 2011 Census of Population
批准号:
ES/L013800/1
负责人:
Paul Longley
金额:
$17.39万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --
中文摘要
先前在伦敦大学学院进行的研究表明,一个名字往往提供了一个开放的和可访问的声明,其承载者的文化,种族和语言特征(例如Mateos等2011)。父母对前名的选择可能会进一步揭示这些特征,而不断变化的时尚往往使名字成为年龄和其他地理和社会特征的有效指标。这些信息已被用于制定姓名的工作分类,并已成功地用于为审计目的补充不完整的数据记录-例如,在衡量不同族裔群体的国民保健服务预防性保健举措的成功程度方面。然而,这些分类是使用不完整的地址登记簿(如公开版的选民名册)和电话簿开发的。迄今为止,在这种研究中使用的数据源存在许多缺点,限制了所产生的分类在应用于新数据集时的有用性:1。这些分类所依据的数据来源提供了不完整的、可能有偏见的一般人口数据。例如,公共选民登记册不包括(年轻人和移民)非选民或(隐私敏感)“选择退出”的个人,公共电话簿的覆盖面不够普遍,名字也很少。2.对名字年龄分布的商业分类通常限于16岁以上的年龄组,补充国家统计局出生姓名数据(例如www.ons.gov.uk/ons/rel/vsob1/baby-names--england-and-wales/2012/stb-baby-names-2012.html)容易出错,因为年幼的儿童可能移居国外,移民可能会携带年幼的儿童。因此,这些来源不允许在任何特定时刻对居住在英国的人口进行全面的快照。虽然“众包”验证是可能的(例如www.onomap.org),但没有比较预测和客观(例如年龄)或自我分配(例如种族)特征的综合方法。4.很少有人关注对“难以接触”的群体进行分类的努力,例如加勒比人,他们的种族可能只能通过姓氏配对之间的微妙联系来确定。5.聚类过程在很大程度上是无空间的,在很大程度上是因为地理覆盖的不均匀性和缺乏高度粒度的位置信息。这项研究将通过使用最好的辅助数据集来解决这些缺点,以开发丰富的分类并进行敏感性分析,以改进和改进其在英国的普遍应用。将利用个人层面的普查数据,通过扩展Mateos等人(2011年)的方法,将名和姓按文化、族裔和语言群体进行分类。至关重要的是,这一分类的结果将首次与关于种族、民族认同、出生国、第一和第二语言以及国籍的个人和家庭数据进行比较。这将使调查分类中明显错误的原因成为可能,并查明这些错误集中在哪些小的地理区域。将通过一个反复的程序,利用安全的在线设施,根据这些结果改进分类,将住户和个人分类结果与人口普查身份计量方面的自我分配进行比较,将有可能使分类工具对文化同化指标敏感,无论是通过异族婚姻还是居住时间,从地方到国家的范围。“姓氏地区”还将用于为2011年ONS产出领域分类添加区域背景。
英文摘要
Previous research conducted at UCL has demonstrated that a name very often provides an open and accessible statement of the cultural, ethnic and linguistic characteristics of its bearer (e.g. Mateos et al 2011). Additional light may be shed upon these characteristics by parental choice of fore-(given) name, while changing fashions often render forenames a valid indicator of age and other geographic and social characteristics. This information has been used to develop working classifications of names, and they have been successfully used to augment incomplete data records for audit purposes - for example in gauging the success of NHS preventive care initiatives across different ethnic groups. However, these classifications have been developed using incomplete address registers (such as the public version of the Electoral Roll) and telephone directories.There are a number of shortcomings to the data sources hitherto used in this kind of research that limit the usefulness of the resulting classifications when applied to new datasets:1. The data sources underlying the classifications provide incomplete and probably biased representations of the population-at-large. For example, public electoral registers do not include (young and immigrant) non-voters or (privacy sensitive) 'opt out' individuals, and public telephone directories provide less than universal coverage and few given names. 2. Commercial classifications of the age profiles of given names are typically restricted to the 16+ age cohorts, and supplementation with ONS birth name data (e.g. www.ons.gov.uk/ons/rel/vsob1/baby-names--england-and-wales/2012/stb-baby-names-2012.html) is error prone because young children may move abroad and immigrants may bring young children with them. Thus these sources do not allow an inclusive snapshot of the population resident in the UK at any specific moment in time.3. Whilst 'crowd sourced' validation is possible (e.g. www.onomap.org), there is no comprehensive means of comparing predicted and objective (e.g. age) or self-assigned (e.g. ethnicity) characteristics. 4. Little focus has been developed upon refining attempts to classify 'hard to reach' groups, such as Caribbeans, whose ethnicity can likely only be ascertained through subtle associations between forename-surname pairings. 5. The clustering procedure has been largely aspatial, in significant part because of unevenness of geographical coverage and the absence of highly granular location information. This research will address these shortcomings through use of the best available secondary dataset for developing an enriched classification and conducting sensitivity analysis to refine and improve its universal application across the UK. Individual level Census data will be used in order to develop a classification of given and family names into cultural, ethnic and linguistic groups, by extending the methodology of Mateos et al (2011). Crucially, and for the first time, the results of this classification will be compared to individual and household data on Ethnicity, National Identity, Country of Birth, First and Second Language Spoken and Nationality. This will make it possible to investigate the causes of apparent errors in the classification, and to identify the small geographic areas in which they are concentrated. Through an iterative procedure, secure online facilities will be used to improve the classification in the light of these results.Comparison of household and individual classification results with self assignments in terms of Census measures of identity will make it possible to make the classification tool sensitive to indicators of cultural assimilation, whether through inter-marriage or duration of residence, at scales from the local to the national.The 'surname regions' will also be used to add regional context to the 2011 ONS Output Area Classification.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
The Routledge Handbook of Census Resources, Methods and Applications: Unlocking the UK 2011 Census
劳特利奇人口普查资源、方法和应用手册:解锁英国 2011 年人口普查
DOI:
--
发表时间:
2016
期刊:
影响因子:
--
作者:
[Gale C]
通讯作者:
Gale C
From Data to Narratives: Scrutinising the Spatial Dimensions of Social and Cultural Phenomena Through Lenses of Interactive Web Mapping
从数据到叙述:通过交互式网络映射镜头审查社会和文化现象的空间维度
DOI:
10.1007/s41651-022-00117-x
发表时间:
2022-06-16
期刊:
Journal of Geovisualization and Spatial Analysis
影响因子:
4
作者:
[]
通讯作者:
DOI:
10.1177/2399808317710132
发表时间:
2017
期刊:
Urban Analytics and City Science
影响因子:
--
作者:
[Harris R]
通讯作者:
Harris R
DOI:
10.1371/journal.pone.0201774
发表时间:
2018
期刊:
PloS one
影响因子:
3.7
作者:
[Kandt J, Longley PA]
通讯作者:
Longley PA
DOI:
10.1002/bse.2050
发表时间:
2018-11-01
期刊:
BUSINESS STRATEGY AND THE ENVIRONMENT
影响因子:
13.4
作者:
[Chintakayala, Phani Kumar, Young, William, Morris, Michelle A.]
通讯作者:
Morris, Michelle A.
共 8 条
Retail Business Datasafe
-
批准号:ES/L011840/1
-
项目类别:Research Grant
-
资助金额:$1510.32万
-
财政年份:2014
-
负责人:Paul Longley
-
依托单位:
The Uncertainty of Identity: Linking Spatiotemporal Information Between Virtual and Real Worlds
-
批准号:EP/J005266/1
-
项目类别:Research Grant
-
资助金额:$155.22万
-
财政年份:2011
-
负责人:Paul Longley
-
依托单位:
海外基金