CNIPA.AI
Back to Search
Dossier

METHOD AND APPARATUS FOR PERFORMING SPEECH ENHANCEMENT, STORAGE MEDIUM, DEVICE, AND PRODUCT

InventionPending
1views
19Claims · 3 independent
§ Ⅰ

Dossier Overview

Applicant

TENCENT TECHONOLOGY (SHENZHEN) COMPANY LIMITED

Inventor

Weixin ZHU; Wei RAO; Yannan WANG; Yifeng HU; Defu SHI; Chenli WAN; Gaoxiong YI

IPC Classification

G10L 17/20G10L 17/2G10L 17/4G10L 17/6G10L 17/18G10L 21/2

CPC Classification

G10L17/20G10L17/2G10L17/4G10L17/6G10L17/18G10L21/2

A speech enhancement method, apparatus, and computer-readable storage medium for training neural networks to enhance speech quality. The method obtains a training set containing training samples, each comprising a sample reference speech, a sample comparison speech from the same sound-producing object, and a mixed speech combining interfering human voice, ambient noise, and the sample comparison speech. Sample voiceprint vectors are extracted from reference speech and sample audio features from mixed speech. A speech enhancement network processes these inputs to output predicted audio features, which are compared against comparison audio features to determine training loss values. The network's weight parameters are iteratively updated based on these loss values until training completion, enabling effective speech enhancement through voiceprint-guided processing.

Original (Chinese)

A speech enhancement method, apparatus, and computer-readable storage medium for training neural networks to enhance speech quality. The method obtains a training set containing training samples, each comprising a sample reference speech, a sample comparison speech from the same sound-producing object, and a mixed speech combining interfering human voice, ambient noise, and the sample comparison speech. Sample voiceprint vectors are extracted from reference speech and sample audio features from mixed speech. A speech enhancement network processes these inputs to output predicted audio features, which are compared against comparison audio features to determine training loss values. The network's weight parameters are iteratively updated based on these loss values until training completion, enabling effective speech enhancement through voiceprint-guided processing.

External Resources