HIGH PERFORMANCE EXECUTION OF STATE SPACE MODELS ON NEURAL NETWORK ACCELERATORS
개요
발명자
Arnab Raha; Arghadip Das; Soumendu Kumar Ghosh; Shamik Kundu; Deepak Abraham Mathaikutty
IPC 분류
CPC 분류
State space model (SSM) neural network operations can be executed efficiently on neural network accelerators by mapping sequential aggregation operations to data-parallel hardware. For cumulative sum operations, the neural network accelerator can perform matrix-to-matrix multiplication with a lower-triangular mask to achieve the same result. For reduce sum operations, the neural network accelerator can perform matrix-to-vector multiplication with a vector mask to achieve the same result. These mappings exploit the parallelism in the neural network accelerators, reduce memory traffic, and leverage sparsity compression and compute skipping for efficiency. Additionally, activation functions can be accelerated using programmable look-up tables during the drain phase. The approach achieves significant latency and energy improvements without hardware changes, enabling high performance deployment of SSM-based models on resource-constrained neural network accelerators.
원문 (중국어)
State space model (SSM) neural network operations can be executed efficiently on neural network accelerators by mapping sequential aggregation operations to data-parallel hardware. For cumulative sum operations, the neural network accelerator can perform matrix-to-matrix multiplication with a lower-triangular mask to achieve the same result. For reduce sum operations, the neural network accelerator can perform matrix-to-vector multiplication with a vector mask to achieve the same result. These mappings exploit the parallelism in the neural network accelerators, reduce memory traffic, and leverage sparsity compression and compute skipping for efficiency. Additionally, activation functions can be accelerated using programmable look-up tables during the drain phase. The approach achieves significant latency and energy improvements without hardware changes, enabling high performance deployment of SSM-based models on resource-constrained neural network accelerators.