SPEECH ENHANCEMENT MODEL TRAINING METHOD AND APPARATUS, DEVICE, MEDIUM, AND PROGRAM PRODUCT
案件概要
出願人
Tencent Technology (Shenzhen) Company Limited
発明者
Feng BAO
IPC分類
CPC分類
This present disclosure relates to a speech enhancement model training method and apparatus, an electronic device, and a storage medium. The method includes: extracting a first audio feature of a to-be-enhanced speech signal through an input layer in each instance of iterative training of an initial speech enhancement model; performing frequency band compression on the first audio feature through a frequency band compression layer, to obtain a dimensionality-reduced second audio feature; performing, through a feature mapping layer, feature mapping on the second audio feature by using a cyclic iteration manner, to obtain a third audio feature, a quantity of output channels of the feature mapping layer increasing progressively in a cyclic iteration process; and inputting the third audio feature to an output layer, to obtain estimated gain information, and performing parameter adjustment on the initial speech enhancement model with reference to true gain information.
原文(中国語)
This present disclosure relates to a speech enhancement model training method and apparatus, an electronic device, and a storage medium. The method includes: extracting a first audio feature of a to-be-enhanced speech signal through an input layer in each instance of iterative training of an initial speech enhancement model; performing frequency band compression on the first audio feature through a frequency band compression layer, to obtain a dimensionality-reduced second audio feature; performing, through a feature mapping layer, feature mapping on the second audio feature by using a cyclic iteration manner, to obtain a third audio feature, a quantity of output channels of the feature mapping layer increasing progressively in a cyclic iteration process; and inputting the third audio feature to an output layer, to obtain estimated gain information, and performing parameter adjustment on the initial speech enhancement model with reference to true gain information.
外部リソース