CNIPA.AI
검색으로 돌아가기
기록

HIGH PERFORMANCE EXECUTION OF STATE SPACE MODELS ON NEURAL NETWORK ACCELERATORS

발명심사 중
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Arnab Raha; Arghadip Das; Soumendu Kumar Ghosh; Shamik Kundu; Deepak Abraham Mathaikutty

IPC 분류

G6F 17/16G6N 3/42

CPC 분류

G6F17/16G6N3/42

State space model (SSM) neural network operations can be executed efficiently on neural network accelerators by mapping sequential aggregation operations to data-parallel hardware. For cumulative sum operations, the neural network accelerator can perform matrix-to-matrix multiplication with a lower-triangular mask to achieve the same result. For reduce sum operations, the neural network accelerator can perform matrix-to-vector multiplication with a vector mask to achieve the same result. These mappings exploit the parallelism in the neural network accelerators, reduce memory traffic, and leverage sparsity compression and compute skipping for efficiency. Additionally, activation functions can be accelerated using programmable look-up tables during the drain phase. The approach achieves significant latency and energy improvements without hardware changes, enabling high performance deployment of SSM-based models on resource-constrained neural network accelerators.

원문 (중국어)

State space model (SSM) neural network operations can be executed efficiently on neural network accelerators by mapping sequential aggregation operations to data-parallel hardware. For cumulative sum operations, the neural network accelerator can perform matrix-to-matrix multiplication with a lower-triangular mask to achieve the same result. For reduce sum operations, the neural network accelerator can perform matrix-to-vector multiplication with a vector mask to achieve the same result. These mappings exploit the parallelism in the neural network accelerators, reduce memory traffic, and leverage sparsity compression and compute skipping for efficiency. Additionally, activation functions can be accelerated using programmable look-up tables during the drain phase. The approach achieves significant latency and energy improvements without hardware changes, enabling high performance deployment of SSM-based models on resource-constrained neural network accelerators.