Abstract:
Nonvolatile AI-edge processors based on near-memory-compute (nvNMC) enable energy-efficient multiply-and-accumulate (MAC) operations with short wakeup latency for edge inference operations. Lossless compression is required for floating-point (FP) neural network (NN) models under on-chip memory capacity constraints; however, this imposes several challenges: 1) data decompression overhead due to lossless FP weight encoding; 2) …