Visualisation Tools for Understanding Big Data

Visualisation Tools for Understanding Big Data
复制标题

DOI:
10.1068/b3903ed
复制
发表时间:
2012-06
期刊:
Environment and Planning B: Planning and Design
影响因子:
--
通讯作者:
J. Cheshire;M. Batty
J. Cheshire;M. Batty
中科院分区:
其他
文献类型:
--
作者:
J. Cheshire;M. Batty

文献摘要

被引文献

相似文献

用于理解大数据的可视化工具 (1) 近年来,在技术不断进步的推动下,“大数据”可视化的创建和发布呈爆炸式增长。在上一篇社论(Batty,2012)中,我们写到了智慧城市的崛起如何通过其常规仪器产生了一代巨大的实时数据集(大数据),这些数据集有可能为我们提供全新的信息,以揭示城市在极短的时间内的小规模运作情况。然而,理解这些数据仍然是一个重大挑战。不久前,从如此大的数据集中绘制现象分布图甚至在分析开始之前就需要数周的数据准备时间。现在每天发布的数据量超过了上一代人典型学术生涯中可以收集到的数据量。事实上,知情意见表明,世界信息(数据)每两年翻一番,去年(2011 年)收集了 1.8 ZB,1 ZB 是 2 的 70 次方,或 10 的 21 次方。事实上,很难想象这个数字,更不用说数据了。 Catone (2011) 表示,它相当于 575 亿部 32 GB iPad 的存储空间,但无论如何类比,这个数字太大了,难以理解。不用说,面对这种扩散,尽管理论上“仍然”可能将这些信息长期存储在数字档案中,但大部分信息都将丢失。私营公司首次收集比中央政府更多的个人信息,推动了数据生产的转变。例如,麦肯锡全球研究所 (2011) 报告称,美国 17 个商业部门中有 15 个现在每家公司平均拥有的数据比美国国会图书馆还多。这种转变在可视化背景下意义重大,因为数据的质量,特别是其代表性,并不一定会随着数量的增加而增加,并且对许多人来说将是一个较低的优先级,特别是与国家人口普查机构相比。这是韩国国家统计局寻求逐步取消十年一次的人口普查的一个主要问题。也就是说,大多数大数据集最能代表城市人口,因此可以为城市复杂性的广泛研究提供很多帮助。它们也越来越频繁地具有时间性,从而为空前规模的流动分析提供了潜力(Batty 和 Cheshire,2011)。数据越来越多地涉及交互和关系、网络和连接,毫不奇怪,当前可视化的前沿是可视化网络。随着城市寻求变得更加智能以及其中的多个基础设施变得更好连接,大数据集的时间元素扩展到越来越多的实时源。例如,在伦敦,从泰晤士河当前的深度到伦敦地铁在网络上的位置以及公交车站的等待时间,一切都存在实时信息。可以通过时间表信息(以帮助计算延误)、乘客流信息和大量社会经济数据集将上下文添加到这些源中。以表格形式,这些数据可以轻松扩展到数十亿行,并且需要数百甚至数千千兆字节的存储空间。那么,在这个不断发展的“大数据”领域,作为研究人员,我们如何做出贡献?可视化可以提供什么?虽然我们不能在这里显示,但可以引导读者查看其显示http://mappinglondon.co.uk/2012/04/17/mapped-every-bus-trip-in-london/,伦敦的公交网络是
Visualisation tools for understanding big data (1) In recent years, fuelled by continuing technological advances, there has been an explosion in the creation and publication of visualisations of what has come to be called ‘big data’. In the last editorial (Batty, 2012), we wrote about how the rise of the smart city through its routine instrumentation is leading to a generation of enormous real-time datasets—big data—that have the potential for providing us with entirely new information to reveal the functioning of cities at ne scale and over very short time periods. nderstanding such data, however, remains a major challenge. Not so long ago, mapping the distributions of phenomena from such large datasets required weeks of data preparation even before its analysis could begin. Now the volume of data released each day exceeds anything that could be collected in the typical academic lifetime of a generation ago. In fact, informed opinion suggests the world’s information (data) is doubling every two years and last year (2011) 1.8 zettabytes was collected, a zettabyte being 2 to the 70th power, or 10 to the 21st power. In fact it is hard to visualise this number, never mind the data. Catone (2011) suggests that it is equivalent to the storage in 57.5 billion 32 GB iPads but, whatever the analogy, the number is too big to comprehend. Needless to say in the face of such proliferation, most of this information will be lost despite the possibility that its long-term storage in digital archives is ‘still’ theoretically possible. Private companies that, for the rst time, collect more personal information than central government have fuelled this transformation in data production. For example, The McKinsey Global Institute (2011) reports that fteen out of seventeen business sectors in the nited States now hold more data on average per company than the Library of Congress. This shift is signi cant in the context of visualisation as the quality of data, especially with regard to its representativeness, does not necessarily increase with volume and will be a lower priority for many, especially in comparison with national census agencies. This is a major issue for the K’s f ce for National Statistics as it seeks to phase out the decennial Population Census. That said, the majority of big datasets are most representative of urban populations and so have much to offer wide-ranging studies of urban complexity. Increasingly, they are also frequently temporal, thus offering the potential for the analysis of ows on an unprecedented scale (Batty and Cheshire, 2011). Data is increasingly about interactions and relations, about networks and connections and it is little surprise that the current cutting edge of visualisation is in visualising networks. The temporal element of big datasets extends to the increasing number of real-time feeds as cities seek to become smarter and the multiple infrastructures within them become better connected. In London, for example, real-time feeds exist for everything from the current depth of the iver Thames through to the position of London nderground trains on the network and waiting times at bus stops. Context can be added to these feeds through timetable information (to help calculate delays), passenger ow information, and a wealth of socioeconomic datasets. In tabulated form these data can easily extend to billions of rows and require hundreds, if not thousands, of gigabytes of storage space. So, in this evolving ‘big data’ landscape how can we, as researchers, contribute and what does visualisation have to offer? Although we cannot show it here, but instead direct readers to its display at http:// mappinglondon.co.uk/2012/04/17/mapped-every-bus-trip-in-london/, London’s bus network is