CNIPA.AI
返回搜索
档案

SYSTEM AND METHOD FOR NEURAL NETWORK MULTILINGUAL SPEECH RECOGNITION

发明专利审中
2浏览
20权利要求 · 3 独立
§ Ⅰ

卷宗概要

发明人

Purvi AGRAWAL; Vikas JOSHI; Basil ABRAHAM; Tejaswi SEERAM; Rupeshkumar Rasiklal MEHTA

IPC 分类

G10L 15/16G10L 15/G10L 15/6G10L 15/22

CPC 分类

G10L15/16G10L15/5G10L15/63G10L15/22

Systems, methods, and computer-readable storage devices are disclosed for improved recognition of multiple languages in audio data. One method including: receiving a trained split head multilingual neural network model, the trained split head multilingual neural network model including shared acoustic model layers and a plurality of projection layers, each projection layer of the plurality of projection layers corresponding to a language that the trained split head multilingual neural network model recognizes; receiving audio data, the audio data including speech in a plurality of languages in the audio data, the speech in the plurality of languages corresponding the language recognized by a projection layer of the plurality of projection layers of the trained split head multilingual neural network model; and classifying one or more languages of the speech of the audio data using the trained split head multilingual neural network model.

原文(中文)

Systems, methods, and computer-readable storage devices are disclosed for improved recognition of multiple languages in audio data. One method including: receiving a trained split head multilingual neural network model, the trained split head multilingual neural network model including shared acoustic model layers and a plurality of projection layers, each projection layer of the plurality of projection layers corresponding to a language that the trained split head multilingual neural network model recognizes; receiving audio data, the audio data including speech in a plurality of languages in the audio data, the speech in the plurality of languages corresponding the language recognized by a projection layer of the plurality of projection layers of the trained split head multilingual neural network model; and classifying one or more languages of the speech of the audio data using the trained split head multilingual neural network model.