IEEE Journal of Solid-State Circuits – Early Access

PROTEUS: A 40 nm Programmable General-Purpose Digital Compute-In-Memory Accelerator With eNVM and Hierarchical ISA for Versatile Edge AI

PROTEUS: A 40 nm Programmable General-Purpose Digital Compute-In-Memory Accelerator With eNVM and Hierarchical ISA for Versatile Edge AI 150 150

Abstract:

We present PROTEUS, an 18 mm2 programmable general-purpose digital compute-in-memory (GP-DCIM) accelerator integrating 4 Mb resistive random access memory (RRAM) and 2.6 Mb tensor static random access memory (SRAM) with a 32-bit hierarchical DCIM instruction set architecture (ISA). PROTEUS features fine-grained 1-D matrix tiling and a reconfigurable DCIM datapath/pipeline for near-100% memory …

View on IEEE Xplore

MIDAS: An Energy-Efficient Microscaling Quantization-Based Digital CIM Accelerator With Spatio-Temporal Cyclic Alignment for Generative AI

MIDAS: An Energy-Efficient Microscaling Quantization-Based Digital CIM Accelerator With Spatio-Temporal Cyclic Alignment for Generative AI 150 150

Abstract:

Digital compute-in-memory (DCIM) and the microscaling (MX) data format have emerged as promising hardware and algorithmic solutions, respectively, for energy-efficient generative AI (GenAI) processing. However, their combination introduces challenges, including redundant computation from mantissa alignment, excessive exponent handling overhead, and accuracy loss from static approximation schemes that fail to accommodate …

View on IEEE Xplore

A 28-nm FD-SOI CMOS Analog-IMC Core Based on PCM Featuring 8 512 × 512-Weight Layers and 28M Weights×TOPs/W/mm2

A 28-nm FD-SOI CMOS Analog-IMC Core Based on PCM Featuring 8 512 × 512-Weight Layers and 28M Weights×TOPs/W/mm2 150 150

Abstract:

In-memory computing (IMC) hardware accelerators for deep neural networks (DNNs) require storing a massive number of coefficients within a single computing macro to avoid performance degradation in multicore clusters. This aspect, often overlooked by common figures of merit (FoMs), can be effectively addressed by phase-change memory (PCM) technology, thanks to …

View on IEEE Xplore

A Third-Harmonic-Enhanced Triple-Push DCO Utilizing Source-Combining Technique

A Third-Harmonic-Enhanced Triple-Push DCO Utilizing Source-Combining Technique 150 150

Abstract:

This article presents a detailed investigation into optimizing the amplitude and phase of the transistor’s terminal voltages to generate a high 3rd-harmonic current in the millimeter-wave (mm-Wave) frequency. Based on the analysis, the digitally controlled source-combining triple-push (SCTP) oscillator is derived to significantly enhance the 3rd-harmonic current by introducing …

View on IEEE Xplore

ASAP: A 28-nm Transformer Training Accelerator With Alternating Sparsity and Asymmetrical Microscaling Precision

ASAP: A 28-nm Transformer Training Accelerator With Alternating Sparsity and Asymmetrical Microscaling Precision 150 150

Abstract:

This work presents ASAP, a 28-nm transformer-training accelerator that combines N:M structured sparsity with asymmetric microscaling floating-point (MXFP) precision through a unified algorithm–hardware co-design. ASAP introduces a progressive sparsity schedule in which pruned compute resources are reassigned to increase numerical precision for important weights and activations, stabilizing optimization …

View on IEEE Xplore

A 15-MHz, 2.7-mm2 Notch-Free Hybrid Magnetic Current Sensor With Feedforward Ripple Cancellation and ±0.8% Local Gain Non-Uniformity

A 15-MHz, 2.7-mm2 Notch-Free Hybrid Magnetic Current Sensor With Feedforward Ripple Cancellation and ±0.8% Local Gain Non-Uniformity 150 150

Abstract:

This article proposes a hybrid magnetic current sensor achieving a 15-MHz bandwidth within a compact 2.7-mm2 area. To mitigate the pole–zero mismatch inherent in the two-stage integrator topology, a dual-output Gm-C integrator with subtractor-based compensation is proposed, achieving a ±0.8% local gain non-uniformity. A wideband feedforward ripple suppression scheme cancels …

View on IEEE Xplore

A High-Linearity Shunt-Based In-Line Current Sensor With Self-Heating Compensation and 14.4 V 2 MHz PWM Rejection

A High-Linearity Shunt-Based In-Line Current Sensor With Self-Heating Compensation and 14.4 V 2 MHz PWM Rejection 150 150

Abstract:

This article presents a cost-effective, fully integrated shunt-resistor-based in-line current sensor that delivers high linearity and strong PWM rejection, enabling precise current measurement and control in dynamic driving systems like robotics, imaging, and audio. To mitigate the self-heating of the on-chip shunt resistor, which degrades linearity, a location-based thermal compensation …

View on IEEE Xplore

A BEV Perception Transformer Accelerator With Saliency-Driven Image/Point Cloud Fusion and Phase-Linked Dataflow in 28 nm CMOS

A BEV Perception Transformer Accelerator With Saliency-Driven Image/Point Cloud Fusion and Phase-Linked Dataflow in 28 nm CMOS 150 150

Abstract:

Deploying advanced Transformer-based models for real-time, high-accuracy multimodal bird’s-eye-view (BEV) perception in autonomous driving imposes substantial hardware demands. To address this, we propose a low-cost, low-power image/point-cloud fusion Transformer accelerator that supports two modes: high-performance driving and ultra-low-power sentry operation. We first propose a cross-modal saliency evaluation mechanism …

View on IEEE Xplore

A 28-nm PVT Inner-Tracking Time-Domain Compute-In-Memory Macro for Edge-AI Devices

A 28-nm PVT Inner-Tracking Time-Domain Compute-In-Memory Macro for Edge-AI Devices 150 150

Abstract:

This article presents an energy-efficient and process-, voltage-, and temperature (PVT)-robust time-domain (TD) compute-in-memory (CIM) macro for edge artificial intelligence (AI) devices. It features: 1) a PVT inner-tracking (PIT) technique that aligns the PVT responses of TD computation and TD quantization, delivering inherent robustness without incurring extra power or circuit …

View on IEEE Xplore