Multimodal deep learning from satellite and street-level imagery for measuring income, overcrowding, and environmental deprivation in urban areas.

Multimodal deep learning from satellite and street-level imagery for measuring income, overcrowding, and environmental deprivation in urban areas.
复制标题

DOI:
10.1016/j.rse.2021.112339
复制
发表时间:
2021-05
影响因子:
13.5
通讯作者:
Ezzati M
Ezzati M
中科院分区:
工程技术1区
文献类型:
--
作者:
Suel E;Bhatt S;Brauer M;Flaxman S;Ezzati M

文献摘要

参考文献

被引文献

相似文献

以大规模和低成本收集的数据(例如卫星和街道图像)有可能大大提高城市不平等测量的分辨率、空间覆盖范围和时间频率。对于给定的地理区域,通常可以获得来自不同来源的多种类型的数据。然而,由于方法上的困难,大多数研究在进行测量时使用单一类型的输入数据。我们提出了两种基于深度学习的方法来联合利用卫星和街道图像来测量城市不平等。我们以伦敦为例,对三个选定的产出进行了案例研究,每个产出都以十分位数来衡量:收入、过度拥挤和环境剥夺。我们使用平均绝对误差(MAE)比较了我们提出的多模态模型和相应的单模态模型的性能。首先,将卫星瓦片附加到街道图像中,以增强可获得街道图像的位置的预测,从而将收入、拥挤程度和生活环境的十分位数单位的准确性提高20%、10%和9%。第二种方法,据我们所知是新颖的,它使用U-Net架构以高空间分辨率对城市中的所有网格单元进行预测(例如,在我们的实验中,对伦敦的3米× 3米像素进行预测)。它可以利用全市范围内可用的卫星图像,以及来自街道级图像的更稀疏的信息,从而将准确性提高6%、10%和11%。我们还展示了两种方法的预测图示例,以直观地突出性能差异。我们的模型利用了街道和卫星图像的信息。提出的多模态测量方法优于单模态测量方法。该模型可以在训练和预测过程中处理缺失的数据。多模式框架可以纳入其他模式(例如航空图像)。应用程序可以扩展到不同的结果。
Data collected at large scale and low cost (e.g. satellite and street level imagery) have the potential to substantially improve resolution, spatial coverage, and temporal frequency of measurement of urban inequalities. Multiple types of data from different sources are often available for a given geographic area. Yet, most studies utilize a single type of input data when making measurements due to methodological difficulties in their joint use. We propose two deep learning-based methods for jointly utilizing satellite and street level imagery for measuring urban inequalities. We use London as a case study for three selected outputs, each measured in decile classes: income, overcrowding, and environmental deprivation. We compare the performances of our proposed multimodal models to corresponding unimodal ones using mean absolute error (MAE). First, satellite tiles are appended to street level imagery to enhance predictions at locations where street images are available leading to improvements in accuracy by 20, 10, and 9% in units of decile classes for income, overcrowding, and living environment. The second approach, novel to the best of our knowledge, uses a U-Net architecture to make predictions for all grid cells in a city at high spatial resolution (e.g. for 3 m × 3 m pixels in London in our experiments). It can utilize city wide availability of satellite images as well as more sparse information from street-level images where they are available leading to improvements in accuracy by 6, 10, and 11%. We also show examples of prediction maps from both approaches to visually highlight performance differences. Our model utilizes information from street-level and satellite images. Proposed multimodal measurement approaches outperform unimodal ones. The model can deal with missing data during training and predictions. Multimodal frameworks can incorporate additional modalities (e.g. aerial images). Applications can be expanded to different outcomes.
DOI: 10.1021/acs.est.7b00891
发表时间: 2017-06-20
影响因子: 11.4
作者:
Apte, Joshua S.;Messier, Kyle P.;Hamburg, Steven P.
通讯作者: Hamburg, Steven P.
DOI: 10.1073/pnas.1700035114
发表时间: 2017-12-12
影响因子: 11.1
作者:
Gebru T;Krause J;Wang Y;Chen D;Deng J;Aiden EL;Fei-Fei L
通讯作者: Fei-Fei L
DOI: 10.1038/s41370-018-0017-1
发表时间: 2019-06-01
影响因子: 4.5
作者:
Larkin, Andrew;Hystad, Perry
通讯作者: Hystad, Perry
DOI: 10.3390/rs12020329
发表时间: 2020-01-01
期刊: REMOTE SENSING
影响因子: 5
作者:
Barbierato, Elena;Bernetti, Iacopo;Saragosa, Claudio
通讯作者: Saragosa, Claudio
DOI: 10.1016/j.fcr.2012.08.008
发表时间: 2013-03-01
影响因子: 5.8
作者:
Lobell, David B.
通讯作者: Lobell, David B.