基于随机森林优化模型的高分辨率人口密度研究
批准号:
42071167
项目类别:
面上项目
资助金额:
56.0 万元
负责人:
刘劲松
依托单位:
学科分类:
人文地理
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
刘劲松
中文摘要
人口密度图是揭示人地关系的重要基础数据。基于随机森林的人口密度模型显著改善了人口密度图的质量,但存在区群谬误、影响因子遴选欠缜密、混淆人口分布规律等问题。本研究以石家庄为实验区,以河北省为模型应用区,依据自然区划、地貌区划和城乡分布,编制综合禀赋分区图;利用村常住人口数据,编制最小粒度人口密度图;从自然、经济和制度禀赋等维度,遴选人口密度影响因子;以公顷网格为采样单元,分区开展随机抽样,克服了样本采集单元面积显著大于输出单元面积(区群谬误)的问题,并避免影响因子数据产生可塑性面积单元问题,改善了模型输入数据的质量;借助随机森林模型,开展分组对照实验,分区筛选人口密度影响因子,分区推敲随机采样规模,分区编制人口密度图,避免在人口密度图中混淆人口分布法则。本研究将改善人口密度模型的再测信度和结构效度,推动人口密度影响机制、人口分布规律等基础理论研究,为编制全球高分辨率人口密度图提供中国方案。
英文摘要
The population density map is an important basic data to reveal the Man-land relationship. The population density algorithm based on the random forest model significantly improves the quality of the population density map, but there are problems such as ecological fallacy, inadequate selection of influence factors, and confusion of population distribution rules. This study uses Shijiazhuang as the experimental area and Hebei Province as the model application area. Based on natural zoning, landform zoning, and urban-rural distribution, a comprehensive endowment zoning map is prepared. Using the village resident population data, the minimum grain population density map is prepared. It selects the population density influencing factors from the dimensions of natural, economic, and institutional endowment; It uses the hectare grid as the sampling unit and conduct random sampling by area, which overcomes the problem that the area of the sample collection unit is significantly larger than the output unit area (ecological fallacy), and avoids modifiable areal unit problem generated by the impact factor data, which improves the quality of model input data; With the help of random forest model, grouped control experiments are carried out, it screens population density impact factors by district, consider random sampling size by district, compile population density map by district, which overall avoids confusing population distribution rules in the population density map. This study will improve the test-retest reliability and structural validity of the population density model, promote basic theoretical research on population density impact mechanisms, and population distribution laws, and thus provide a Chinese solution for the preparation of a global high-resolution population density map.
在整合人口、资源和环境数据时,联合国可持续发展目标认为,人口密度模型的再测信度和准则效度亟待改进。为解决可塑性面积单位问题、区群谬误问题、混淆人口分布规律问题、无法开展再测信度检验问题对人口密度模型的困扰,项目组在石家庄、河北省、太行山等地先后开展了“自上而下人口密度随机森林分解算法”和“自下而上人口密度随机森林估计算法”的优化实验,在青藏高原藏南地区开展了“自上而下人口密度随机森林估计算法”的优化实验;揭示了石家庄和河北省人口分布规律和影响机制;揭示了2000年以来太行山人口分布时空演变过程和影响机制,探讨了太行山人口结构格局与人口规模增减过程的时空耦合关系;开展了1982年以来青藏高原人口分布格局-过程-机制研究。取得了“基于聚落的人口统计数据空间分解算法、优化人口密度随机森林模型的掩膜系统、人口密度随机森林模型优化实验研究、基于人口密度随机森林优化模型的自下而上人口估计算法、基于人口密度随机森林优化模型的自上而下人口估计算法、河北省人口分布规律及影响机制”等系列研究成果。制备了2007年石家庄人口密度栅格数据集(100m×100m);2014年河北省人口密度栅格数据集(100m×100m);2000年、2010年、2020年太行山人口密度栅格数据集(1km×1km);1982年、1990年、2000年、2010年、2020年青藏高原人口密度栅格数据集(1km×1km)。形成了“标签保真、掩膜控制、分区建模、分层采样、优选样本、遴选因子、加权输出、分区密度制图”的人口密度随机森林模型的整体优化方案,显著改善了人口密度栅格数据集的再测信度和准则效度,为揭示区域人口分布规律和影响机制、为分性别分年龄别凝练区域人口分布模式和影响机制、为精准测算任意自然地理单元人口规模及增减分化过程,提供了统一的技术框架;为编制全球、“一带一路”和全国人口密度栅格数据集,提供了中国方案;为开展多维人口数据空间化(人口规模、自然结构、社会经济结构、流入流出强度),生产行政和自然单元高度自洽的人口数据集,提供关键技术。
泥河湾盆地衰落农村发展模式与演化规律研究
-
批准号:41671138
-
项目类别:面上项目
-
资助金额:75.0万元
-
批准年份:2016
-
负责人:刘劲松
-
依托单位:
人口流动背景下的河北省人口发展功能区演化趋势研究
-
批准号:41240004
-
项目类别:专项基金项目
-
资助金额:20.0万元
-
批准年份:2012
-
负责人:刘劲松
-
依托单位:
基于可塑性面积单元问题的人口密度研究
-
批准号:40871073
-
项目类别:面上项目
-
资助金额:44.0万元
-
批准年份:2008
-
负责人:刘劲松
-
依托单位:
国内基金
海外基金