Artificial intelligence (AI)

A 28-nm FD-SOI CMOS Analog-IMC Core Based on PCM Featuring 8 512 × 512-Weight Layers and 28M Weights×TOPs/W/mm2

A 28-nm FD-SOI CMOS Analog-IMC Core Based on PCM Featuring 8 512 × 512-Weight Layers and 28M Weights×TOPs/W/mm2 150 150

Abstract:

In-memory computing (IMC) hardware accelerators for deep neural networks (DNNs) require storing a massive number of coefficients within a single computing macro to avoid performance degradation in multicore clusters. This aspect, often overlooked by common figures of merit (FoMs), can be effectively addressed by phase-change memory (PCM) technology, thanks to …

View on IEEE Xplore

A Two-Step Time-Domain Logic-Compatible Embedded-Flash Computing-in-Memory Macro for On-Device AI

A Two-Step Time-Domain Logic-Compatible Embedded-Flash Computing-in-Memory Macro for On-Device AI 150 150

Abstract:

Computing in memory (CIM) has emerged as a promising architecture that can significantly reduce power consumption and latency in edge artificial intelligence (AI) devices. For battery-constrained edge devices, non-volatile memories (NVMs) are ideal due to their ability to retain neural network parameters even when powered down, ensuring rapid response time …

View on IEEE Xplore

A 28-nm Sign-Bit-Embedded Bit-Parallel SRAM Compute-in-Memory Macro With Stage-Wise-Enabled Transition-Counting-Lines and Accumulators for Edge-AI Devices

A 28-nm Sign-Bit-Embedded Bit-Parallel SRAM Compute-in-Memory Macro With Stage-Wise-Enabled Transition-Counting-Lines and Accumulators for Edge-AI Devices 150 150

Abstract:

Computing-in-memory (CIM) architectures are promising for edge-AI devices, as they mitigate the von Neumann bottleneck and reduce data movement. Bit-parallel CIM macros with direct memory mapping and low toggle rate are well-suited for computing systems. However, severe routing congestion, inefficient signed computation, and unnecessary toggling of intermediate signals in local …

View on IEEE Xplore

A Nonvolatile AI-Edge Processor With Lossless-Compressed-Computing STT-MRAM Near-Memory-Compute Macro Using Dynamic Floating-/Fixed-Point Accumulation

A Nonvolatile AI-Edge Processor With Lossless-Compressed-Computing STT-MRAM Near-Memory-Compute Macro Using Dynamic Floating-/Fixed-Point Accumulation 150 150

Abstract:

Nonvolatile AI-edge processors based on near-memory-compute (nvNMC) enable energy-efficient multiply-and-accumulate (MAC) operations with short wakeup latency for edge inference operations. Lossless compression is required for floating-point (FP) neural network (NN) models under on-chip memory capacity constraints; however, this imposes several challenges: 1) data decompression overhead due to lossless FP weight encoding; 2) …

View on IEEE Xplore

OISMA: On-the-Fly In-Memory Stochastic Multiplication Architecture for Approximate Matrix Multiplication

OISMA: On-the-Fly In-Memory Stochastic Multiplication Architecture for Approximate Matrix Multiplication 150 150

Abstract:

Artificial intelligence (AI) models are currently driven by a significant upscaling of their complexity, with massive matrix-multiplication workloads representing the major computational bottleneck. In-memory computing (IMC) architectures are proposed to avoid the von Neumann bottleneck. However, both digital/binary-based and analog IMC architectures suffer from various limitations, which significantly degrade …

View on IEEE Xplore

A 28-nm Digital Transpose SRAM Compute-in-Memory Macro With Accurate/Approximate Dual Mode for Floating-Point Edge Training and Inference

A 28-nm Digital Transpose SRAM Compute-in-Memory Macro With Accurate/Approximate Dual Mode for Floating-Point Edge Training and Inference 150 150

Abstract:

Static random-access memory (SRAM)-based computing-in-memory (CIM) macros have been widely studied to improve the energy efficiency of edge artificial intelligence (AI) inference tasks. However, less attention has been given to AI training, which requires CIM macros to not only perform matrix multiply-accumulate (MAC) operations but also support matrix transposition. …

View on IEEE Xplore

A 14-nm Nonvolatile-Volatile-Fused Compute-In-Memory Macro Based on Logic-Compatible Flash for Plastic Neural Networks

A 14-nm Nonvolatile-Volatile-Fused Compute-In-Memory Macro Based on Logic-Compatible Flash for Plastic Neural Networks 150 150

Abstract:

Designing computing-in-memory (CIM) chips with synaptic plasticity can potentially support energy-efficient on-chip learning in edge devices for rapid local task adaptation. Its silicon implementation is challenging as it requires hybridizing nonvolatile and volatile memory (VM) and customized computational operations. In this work, we propose a plastic CIM (P-CIM) macro featuring: 1) …

View on IEEE Xplore

A Microscaling Multi-Mode Gain-Cell Computing-in-Memory Macro for Advanced AI Edge Device

A Microscaling Multi-Mode Gain-Cell Computing-in-Memory Macro for Advanced AI Edge Device 150 150

Abstract:

The microscaling (MX) format is an emerging data representation that quantizes high-bitwidth floating-point (FP) values into low-bitwidth FP-like values with a shared-scale (SS) exponent. When implemented with computing-in-memory (CIM), MX allows an attractive tradeoff between accuracy and hardware efficiency for specific neural network (NN) workloads. This work presents the first …

View on IEEE Xplore