Computer integrated manufacturing

A Two-Step Time-Domain Logic-Compatible Embedded-Flash Computing-in-Memory Macro for On-Device AI

A Two-Step Time-Domain Logic-Compatible Embedded-Flash Computing-in-Memory Macro for On-Device AI 150 150

Abstract:

Computing in memory (CIM) has emerged as a promising architecture that can significantly reduce power consumption and latency in edge artificial intelligence (AI) devices. For battery-constrained edge devices, non-volatile memories (NVMs) are ideal due to their ability to retain neural network parameters even when powered down, ensuring rapid response time …

View on IEEE Xplore

MTDiff: A Multi-Task Diffusion Accelerator With Bandwidth-Aware Memory Partition and Efficient Digital Compute-in-Memory

MTDiff: A Multi-Task Diffusion Accelerator With Bandwidth-Aware Memory Partition and Efficient Digital Compute-in-Memory 150 150

Abstract:

Initially applied to image synthesis, diffusion models have rapidly expanded into various content generation tasks, such as 3-D scene reconstruction and video generation, due to the superior generative quality. Based on our observations, the imbalanced operational intensities across diffusion model layers lead to suboptimal hardware utilization due to inadequate bandwidth …

View on IEEE Xplore

APM-CIM: An Array Partition Multi-Macro CIM System With Dynamic Sparse Approximation for Neural Network Edge Applications

APM-CIM: An Array Partition Multi-Macro CIM System With Dynamic Sparse Approximation for Neural Network Edge Applications 150 150

Abstract:

Computing-in-memory (CIM) has been widely investigated as a solution to the von Neumann bottleneck to reduce data transfer between memory and processor. However, existing CIM macros for neural network (NN) edge applications still suffer from high latency and energy consumption due to the correlation between data precision and the number …

View on IEEE Xplore

A 28-nm Sign-Bit-Embedded Bit-Parallel SRAM Compute-in-Memory Macro With Stage-Wise-Enabled Transition-Counting-Lines and Accumulators for Edge-AI Devices

A 28-nm Sign-Bit-Embedded Bit-Parallel SRAM Compute-in-Memory Macro With Stage-Wise-Enabled Transition-Counting-Lines and Accumulators for Edge-AI Devices 150 150

Abstract:

Computing-in-memory (CIM) architectures are promising for edge-AI devices, as they mitigate the von Neumann bottleneck and reduce data movement. Bit-parallel CIM macros with direct memory mapping and low toggle rate are well-suited for computing systems. However, severe routing congestion, inefficient signed computation, and unnecessary toggling of intermediate signals in local …

View on IEEE Xplore

A 194.6-TOPS/W Pipelined All Current-Domain Mixed-Signal Compute in Memory in 28-nm CMOS

A 194.6-TOPS/W Pipelined All Current-Domain Mixed-Signal Compute in Memory in 28-nm CMOS 150 150

Abstract:

Mixed-signal CIM (MS-CIM) faces bit-cell nonlinearity, poor linearity at high frequency, and throughput limits. We present a hybrid pipelined current-domain MS-CIM macro featuring bit-cell matched linearization interface (BMLI) and loop-unrolled successive approximation refinement (SAR) ADC fabricated in 28-nm CMOS. A $256{\,}\times {\,}256$ SRAM array with 8-bit inputs, 8-bit weights achieve 10.16-TOPS …

View on IEEE Xplore

A 28-nm PVT Inner-Tracking Time-Domain Compute-In-Memory Macro for Edge-AI Devices

A 28-nm PVT Inner-Tracking Time-Domain Compute-In-Memory Macro for Edge-AI Devices 150 150

Abstract:

This article presents an energy-efficient and process-, voltage-, and temperature (PVT)-robust time-domain (TD) compute-in-memory (CIM) macro for edge artificial intelligence (AI) devices. It features: 1) a PVT inner-tracking (PIT) technique that aligns the PVT responses of TD computation and TD quantization, delivering inherent robustness without incurring extra power or circuit …

View on IEEE Xplore