Data Visualisation: Learning From Big Data
Data Visualisation: Learning From Big Data
批准号:
2640731
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
这个名为“数据可视化:从大数据中学习”的研究项目的目的是回答这个问题:“工程师如何应对当前大数据的挑战,特别是那些关于数据量,速度和高维的挑战,以便他们可以更有效地从模拟和实验中提取有意义的见解?".目前,研究工程师面临的许多困难之一是在生产或具有成本效益的时间范围内存储,发送,恢复和查看大量仿真数据。这个问题,在文献中通常被称为“I/O瓶颈”,只会因为硬件加速的高性能计算的成熟所带来的计算能力的持续进步而加剧。例如,在涡轮机械等研究领域,最先进的直接数值模拟(DNS)现在产生单快照输出,每个输出需要数十千兆字节的存储空间。随着平均计算速度超过当前最佳数据传输速率近一个数量级,使用当前技术和基础设施检查和操作如此大量的数据几乎是不可能的。因此,必须考虑新的方法。该项目的最初目标是研究、开发并成功实施I/O瓶颈问题的软件解决方案,同时关注海量集合数据集和大型奇异数据量应用程序。最突出的方法是将压缩作为一种手段,以更紧凑的形式表示大量的多维仿真或实验数据。神经隐式表示被探索作为一种途径,实现极端的数据压缩,而不显着的信息丢失。一个实际的目标是在基于Web的查看器中实时可视化并最终与DNS模拟数据交互:这是一项以当前功能不易完成的任务。工作还将探索神经表征的训练方案,以及减压过程的各个方面。其他研究目标是调查和应用统计工具,数据科学方法和机器学习(ML)方法来帮助自动生成洞察力。这项工作的目标是降低在进行探索性数据分析(EDA)时对特定领域知识和统计熟练程度的需求。最后,延伸目标还包括探索新的和颠覆性的技术,用于交互式数据可视化。
英文摘要
The aim of this research project entitled "Data Visualisation: Learning From Big Data" is to answer the question: "how can engineers contend with the current challenges of Big Data, particularly those concerning data volume, velocity and high-dimensionality, so they may more effectively extract meaningful insights from their simulations and experiments?". At present, one of the many difficulties faced by research engineers is that of storing, sending, recovering and viewing extreme volumes of simulation data in productive or cost-effective time frames. This issue, often referred to in the literature as the "I/O bottleneck", is only going to be exacerbated by continued advances in compute capability rooted in the maturation of hardware accelerated high-performance computing. For example, in fields of study such as Turbomachinery, cutting edge Direct Numerical Simulations (DNS) now yield single-snapshot outputs each demanding tens of GigaBytes of storage space. With average compute speeds now exceeding current best data transfer rates by nearly an order of magnitude, it is becoming almost impossible to inspect and manipulate such large volumes of data with current techniques and infrastructure. As such, new approaches must be considered.The initial aims of this project are to investigate, develop, and successfully implement software solutions to the I/O bottleneck problem, focussing evenly on massive ensemble dataset and large singular data volume applications. The preeminent approach is to implement compression as a means of representing large multidimensional volumes of simulation or experimental data in more compact forms. Neural implicit representations are to be explored as an avenue for achieving extreme data compression without significant loss of information. A practical target is to visualise and ultimately interact with DNS simulation data in real-time within a web-based viewer: a task that is not easily accomplished with current capabilities. Work will also explore training schemes for neural representations, as well as aspects of the decompression process. Additional research aims are to investigate and apply statistical tools, data science methods, and machine learning (ML) methods to assist in automatic insight generation. The goal of this work will be to lower the demand for both domain-specific knowledge and proficiency in statistics when performing exploratory data analysis (EDA). Finally, stretch goals are to also explore new and disruptive technologies for application in interactive data visualisation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金