Impact Factor
Call For Paper
Volume 12 Issue 09
September 2026
Author(s)
Abstract
CNN Inference On FPGA Platforms Can Incur Substantial Dynamic Power Because Parallel Multiply-accumulate Datapaths May Continue Switching When Sparse Operands Contain Zeros. This Paper Presents A Sparsity-aware Operand-isolated Systolic Array Accelerator For CNN Computation. The Architecture Comprises A Regular 4×4 Array Of 16 Processing Elements, Using 8-bit Operands And 32-bit Accumulation. Each Processing Element Performs Local Zero Detection At Its Multiplier Inputs. When Either Operand Is Zero, The Multiplier Inputs Are Isolated To Zero And The Accumulator Update Is Disabled; Otherwise, Normal Multiply-accumulate Operation Proceeds. The Regular Systolic Dataflow Is Therefore Retained Without A Separate Sparse Scheduler Or Irregular Interconnect. The Design Was Implemented For The Xilinx Artix-7 XC7A35TCPG236-1 FPGA. Post-implementation Results Report 590 LUTs, 213 Flip-flops, And 16 DSP Resources, With Utilizations Of 2.84%, 0.51%, And 17.78%, Respectively. The Reported Timing Period Is 9.524 Ns, Corresponding To Approximately 104.998 MHz. Power Analysis Was Performed Separately At 27 MHz Using Simulation-derived Activity. Total On-chip Power Was 90, 89, 77, And 71 MW For 0%, 50%, 75%, And 100% Sparsity. Dynamic Power Decreased From 22 MW To 3 MW, Corresponding To An 86.36% Reduction. The Results Indicate That Local Operand Isolation And Accumulator-enable Control Can Reduce Dynamic Switching In A Compact Regular Systolic Accelerator.
Keywords
Paper ID
IJSARTV12I8105839
Publication Date
August 30, 2026
Research Area
VLSI