DYNAMIC SPARSITY PATTERNS FOR ATTENTION HEADS
案件概要
出願人
Microsoft Technology Licensing, LLC
発明者
Huiqiang JIANG; Yucheng LI; Chengruidong ZHANG; Qianhui WU; Xufang LUO; Surin Lindsay AHN; Zhenhua HAN; Amir H. ABDI; Dongsheng LI; Chin-Yew LIN; Yuqing YANG; Lili QIU
IPC分類
CPC分類
A computing system including processing circuitry configured to, during a calibration stage, perform a sparsity pattern search on a plurality of attention heads included in one or more transformer layers to select a respective sparsity pattern associated with each of the attention heads. During an inferencing stage, processing circuitry receives an inferencing input. The processing circuitry pre-fills a context based at least in part on the inferencing input. Pre-filling the context includes computing sparse attention scores at each of the attention heads. Computing the sparse attention scores includes masking each of the attention heads using the respective sparsity pattern selected for that attention head during the calibration stage. The processing circuitry computes an inferencing output by performing inferencing starting from the sparse attention scores. The processing circuitry outputs the inferencing output.
原文(中国語)
A computing system including processing circuitry configured to, during a calibration stage, perform a sparsity pattern search on a plurality of attention heads included in one or more transformer layers to select a respective sparsity pattern associated with each of the attention heads. During an inferencing stage, processing circuitry receives an inferencing input. The processing circuitry pre-fills a context based at least in part on the inferencing input. Pre-filling the context includes computing sparse attention scores at each of the attention heads. Computing the sparse attention scores includes masking each of the attention heads using the respective sparsity pattern selected for that attention head during the calibration stage. The processing circuitry computes an inferencing output by performing inferencing starting from the sparse attention scores. The processing circuitry outputs the inferencing output.
外部リソース