Analog content-addressable memory from complementary FeFETs

Analog content-addressable memory from complementary FeFETs
复制标题

来自互补 FeFET 的模拟内容寻址存储器

DOI:
10.1016/j.device.2023.100218
复制
发表时间:
2024
期刊:
Device
影响因子:
--
通讯作者:
Jariwala, Deep
Jariwala, Deep
中科院分区:
--
文献类型:
--
作者:
Liu, Xiwen;Katti, Keshava;He, Yunfei;Jacob, Paul;Richter, Claudia;Schroeder, Uwe;Kurinec, Santosh;Chaudhari, Pratik;Jariwala, Deep

文献摘要

相似文献

尽管最近在用于矩阵乘法的非易失性存储器(NVM)方面取得了进展,但其他关键的数据密集型操作(如并行搜索)仍然在很大程度上被忽视。当前的并行搜索架构,即内容可寻址存储器(CAM),通常使用二进制,这限制了密度和功能。我们提出了一个模拟CAM(ACAM)单元,建立在两个互补的铁电场效应晶体管(FeFET),在模拟域中进行并行搜索超过40个不同的匹配窗口。ACAM不仅提供了比三元CAM(TCAM)更密集3倍的内存架构,而且对于使用Omniglot数据集模拟的少次学习,在相似性搜索方面的推理准确性提高了5%,与基于硅基互补金属氧化物半导体(CMOS)节点。此外,我们在ACAM中演示了对内核回归模型的一步推理,模拟结果表明推理速度比CPU和GPU快1,000倍。
Despite recent advancements in non-volatile memory (NVM) for matrix multiplication, other critical data-intensive operations like parallel search remain largely overlooked. Current parallel search architectures, namely content-addressable memory (CAM), often use binary, which restricts density and functionality. We present an analog CAM (ACAM) cell, built on two complementary ferroelectric field-effect transistors (FeFETs), that performs parallel search in the analog domain with over 40 distinct match windows. ACAM not only offers a projected 3× denser memory architecture than ternary CAM (TCAM) but also yields a 5% increase in inference accuracy on similarity search for few-shot learning simulated with the Omniglot dataset, with an estimated speedup per similarity search of more than 100× when compared to a central processing unit (CPU) and graphics processing unit (GPU) on scaled silicon-based complementary metal-oxide semiconductor (CMOS) nodes. Furthermore, we demonstrate one-step inference on a kernel regression model in ACAM, with simulation results indicating 1,000× faster inference than a CPU and GPU.