Hard real-time core software of the AO RTC COSMIC platform: architecture and performance

Hard real-time core software of the AO RTC COSMIC platform: architecture and performance
复制标题

AO RTC COSMIC平台硬实时核心软件:架构与性能

DOI:
--
复制
发表时间:
2020
期刊:
Astronomical Telescopes + Instrumentation
影响因子:
--
通讯作者:
D. Gratadour
D. Gratadour
中科院分区:
--
文献类型:
--
作者:
F. Ferreira;A. Sevin;J. Bernard;O. Guyon;A. Bertrou;J. Raffard;F. Vidal;E. Gendron;D. Gratadour

文献摘要

被引文献

相似文献

随着即将到来的巨型望远镜,自适应光学(AO)变得比以往任何时候都更加重要,以获得这些望远镜提供的全部潜力。这种AO系统的复杂性正在达到极端的高度,为了建造它们,必须进行破坏性的开发。AO系统的关键组件之一是真实的时间控制器(RTC),它必须计算0.5到几kHz范围内的高频下的斜率和可变形镜(DM)命令。由于RTC中涉及的计算的复杂性随着望远镜的尺寸而增加,因此满足超大望远镜(ELT)级的RTC要求是一项挑战。例如,MICADO SCAO(单共轭自适应光学)系统需要大约1 TMAC/s的RTC才能获得足够的性能。这种复杂性带来了对高性能计算(HPC)技术和标准的需求,例如使用GPU等硬件加速器。最重要的是,构建RTC通常依赖于项目,因为组件和接口从一种仪器到另一种仪器会发生变化。COSMIC平台旨在开发一个通用的AO RTC平台,该平台功能强大,模块化,可供AO社区使用。这一发展是巴黎天文台和澳大利亚国立大学(ANU)与斯巴鲁望远镜合作的共同努力。我们在这里重点介绍该平台的核心硬实时组件的当前状态。H-RTC流水线由业务单元(BU)组成:每个BU是一个独立的进程,负责一个特定的操作,例如矩阵向量乘法(MVM)或质心计算,可以在CPU或GPU上进行。BU在CACAO框架处理的共享内存(SHM)上读取和写入数据。然后,可以通过使用信号量或通过在GPU上等待的忙碌来进行每个BU之间的同步,以确保非常低的抖动。然后可以通过Python接口控制RTC管道。这种架构的一个关键点是,BU与各种SHM的接口被抽象化,因此在可用BU的集合中添加新BU是直接的。这种方法允许高性能,可扩展,模块化和可配置的RTC管道,可以满足任何AO系统配置的需求。性能已经在MICADO SCAO规模的RTC管道上进行了测量,该管道在配备8个Tesla V100 GPU的DGX-1系统上具有约25,000个斜坡和5,000个执行器。所考虑的流水线由两个BU组成:第一个输入原始金字塔WFS图像(由模拟器产生),在其上应用暗和平坦的参考,然后从图像中提取有用的像素。第二BU执行MVM和遵循经典积分器命令法则的命令的积分。BU之间的同步通过GPU忙碌等待BU输入来实现。所获得的性能表明,使用4个GPU时,平均延迟高达235 μ s,抖动为4.4 μ s rms,最大抖动为30 μ s
With the upcoming giant class of telescopes, Adaptive Optics (AO) has become more essential than ever before to get access to the full potential offered by those telescopes. The complexity of such AO systems is reaching extreme heights, and disruptive developments will have to be made in order to build them. One of the critical component of a AO system is the Real Time Controller (RTC) which will have to compute the slopes and the Deformable Mirror (DM) commands at high frequency, in a range of 0.5 to several kHz. Since the complexity of the computations involved in the RTC is increasing with the size of the telescope, fulfilling RTC requirements for Extremely Large Telescope (ELT) class is a challenge. As an example, the MICADO SCAO (Single Conjugate Adaptive Optics) system requires around 1 TMAC/s for the RTC to get sufficient performance. This complexity brings the need for High Performance Computing (HPC) techniques and standards, such as the use of hardware accelerator like GPU. On top of that, building a RTC is often project-dependent as the components and the interfaces change from one instrument to an other. The COSMIC platforms aims at developing a common AO RTC platform which is meant to be powerful, modular and available to the AO community. This development is a joint effort between Observatoire de Paris and the Australian National University (ANU) in collaboration with the Subaru Telescope. We focus here on the current status of the core hard real-time component of this platform. The H-RTC pipeline is composed of Business Units (BU): each BU is an independent process in charge of one particular operation, such as Matrix Vector Multiply (MVM) or centroid computation, that can be made on CPU or on GPU. BUs read and write data on Shared Memory (SHM) handled by the CACAO framework. Synchronization between each BU can then be made either by using semaphore or by busy waiting on the GPU to ensure very low jitter. The RTC pipeline can then be controlled through a Python interface. One of the key point of this architecture is that the interfaces of a BU with the various SHM is abstracted, so adding a new BU in the collection of available ones is straight forward. This approach allows a high performance, scalable, modular and configurable RTC pipeline that could fit the needs of any AO system configuration. Performance has been measured on a MICADO SCAO scale RTC pipeline with around 25,000 slopes by 5,000 actuators on a DGX-1 system equipped with 8 Tesla V100 GPUs. The considered pipeline is composed of two BUs : the first one takes an input the raw pyramid WFS image (produced by simulator), applies on it dark and flat references, and then extract the useful pixel from the image. The second BU performs the MVM and the integration of the commands following a classical integrator command law. Synchronization between the BU is made through GPU busy waiting on the BU inputs. Performance obtained shows a mean latency up to 235 μs using 4 GPUs, with a jitter of 4.4 μs rms and a maximum jitter of 30 μs