CNIPA.AI
返回搜索
档案

METHOD AND APPARATUS FOR PERFORMING SPEECH ENHANCEMENT, STORAGE MEDIUM, DEVICE, AND PRODUCT

发明专利审中
9浏览
19权利要求 · 3 独立
§ Ⅰ

卷宗概要

发明人

Weixin ZHU; Wei RAO; Yannan WANG; Yifeng HU; Defu SHI; Chenli WAN; Gaoxiong YI

IPC 分类

G10L 17/20G10L 17/2G10L 17/4G10L 17/6G10L 17/18G10L 21/2

CPC 分类

G10L17/20G10L17/2G10L17/4G10L17/6G10L17/18G10L21/2

A speech enhancement method, apparatus, and computer-readable storage medium for training neural networks to enhance speech quality. The method obtains a training set containing training samples, each comprising a sample reference speech, a sample comparison speech from the same sound-producing object, and a mixed speech combining interfering human voice, ambient noise, and the sample comparison speech. Sample voiceprint vectors are extracted from reference speech and sample audio features from mixed speech. A speech enhancement network processes these inputs to output predicted audio features, which are compared against comparison audio features to determine training loss values. The network's weight parameters are iteratively updated based on these loss values until training completion, enabling effective speech enhancement through voiceprint-guided processing.

原文(中文)

A speech enhancement method, apparatus, and computer-readable storage medium for training neural networks to enhance speech quality. The method obtains a training set containing training samples, each comprising a sample reference speech, a sample comparison speech from the same sound-producing object, and a mixed speech combining interfering human voice, ambient noise, and the sample comparison speech. Sample voiceprint vectors are extracted from reference speech and sample audio features from mixed speech. A speech enhancement network processes these inputs to output predicted audio features, which are compared against comparison audio features to determine training loss values. The network's weight parameters are iteratively updated based on these loss values until training completion, enabling effective speech enhancement through voiceprint-guided processing.