Optimizing Near-ML MIMO Detector for SDR Baseband on Parallel Programmable Architectures

Optimizing Near-ML MIMO Detector for SDR Baseband on Parallel Programmable Architectures
复制标题

在并行可编程架构上优化 SDR 基带的近 ML MIMO 检测器

DOI:
--
复制
发表时间:
2008
期刊:
Design, Automation and Test in Europe
影响因子:
--
通讯作者:
F. Catthoor
F. Catthoor
中科院分区:
--
文献类型:
--
作者:
Min Li;B. Bougard;Weiyu Xu;D. Novo;L. Perre;F. Catthoor

文献摘要

被引文献

相似文献

近年来,ML和近ML MIMO探测器引起了很多兴趣。但是,几乎所有报告的实现都在ASIC或FPGA中提供。我们的贡献是优化近ML MIMO检测器,用于并行可编程架构,例如具有ILP和DLP功能的架构。在拟议的SSFE(选择性枚举的选择性跨越)中,从设计流的一开始就明确引入了建筑友好性。重要的是,高级算法转换使数据流模式和结构拟合体系结构特征非常好。我们可以在SSFE中使用高规律和确定性的数据流启用丰富的矢量并行性;记忆重新排列,改组和不可预测的动力被精心排除在外。因此,SSFE可以轻松并行并有效地映射到ILP和DLP架构上。此外,为了对并行体系结构微调SSFE,在应用程序级信息的帮助下应用了广泛的预补偿器转换。这些不仅优化了计算操作,还优化了解决生成和记忆访问。实验表明,SSFE对现实生活中的VLIW架构带来了非常有效的资源利用。具体而言,在SSFE中,VLIW上的NOP指令百分比低于1%,甚至比软件涉及FFT所实现的NOP指令要好得多。据我们所知,这是有关近ML MIMO探测器的全面优化用于并行可编程体系结构的首次报道的工作。
ML and near-ML MIMO detectors have attracted a lot of interest in recent years. However, almost all the reported implementations are delivered in ASICs or FPGAs. Our contribution is optimizing the near-ML MIMO detector for parallel programmable architectures, such as those with ILP and DLP features. In the proposed SSFE (selective spanning with fast enumeration), architecture-friendliness is explicitly introduced from the very beginning of the design flow. Importantly, high level algorithmic transformations make the dataflow pattern and structure fit architecture-characteristics very well. We enable abundant vector-parallelism with highly regular and deterministic dataflow in the SSFE; memory rearrangements, shuffling and non-predictable dynamism are all elaborately excluded. Hence, the SSFE can be easily parallelized and efficiently mapped onto ILP and DLP architectures. Furthermore, to fine-tune the SSFE on parallel architectures, extensive pre-compiler transformations are applied with the help of the application-level information. These optimize not only computation-operations but also address-generations and memory-accesses. Experiments show that the SSFE brings very efficient resource-utilizations on real-life VLIW architectures. Specifically, with the SSFE the percentage of NOPs instructions on VLIW is below 1%, even better than that achieved by the software-pipelined FFT. To the best of our knowledge, this is the first reported work about comprehensive optimizations of near-ML MIMO detectors for parallel programmable architectures.