SPEECH ENHANCEMENT MODEL TRAINING METHOD AND APPARATUS, DEVICE, MEDIUM, AND PROGRAM PRODUCT
개요
출원인
Tencent Technology (Shenzhen) Company Limited
발명자
Feng BAO
IPC 분류
CPC 분류
This present disclosure relates to a speech enhancement model training method and apparatus, an electronic device, and a storage medium. The method includes: extracting a first audio feature of a to-be-enhanced speech signal through an input layer in each instance of iterative training of an initial speech enhancement model; performing frequency band compression on the first audio feature through a frequency band compression layer, to obtain a dimensionality-reduced second audio feature; performing, through a feature mapping layer, feature mapping on the second audio feature by using a cyclic iteration manner, to obtain a third audio feature, a quantity of output channels of the feature mapping layer increasing progressively in a cyclic iteration process; and inputting the third audio feature to an output layer, to obtain estimated gain information, and performing parameter adjustment on the initial speech enhancement model with reference to true gain information.
원문 (중국어)
This present disclosure relates to a speech enhancement model training method and apparatus, an electronic device, and a storage medium. The method includes: extracting a first audio feature of a to-be-enhanced speech signal through an input layer in each instance of iterative training of an initial speech enhancement model; performing frequency band compression on the first audio feature through a frequency band compression layer, to obtain a dimensionality-reduced second audio feature; performing, through a feature mapping layer, feature mapping on the second audio feature by using a cyclic iteration manner, to obtain a third audio feature, a quantity of output channels of the feature mapping layer increasing progressively in a cyclic iteration process; and inputting the third audio feature to an output layer, to obtain estimated gain information, and performing parameter adjustment on the initial speech enhancement model with reference to true gain information.