検索に戻る
案件記録

SYSTEM AND METHOD FOR NEURAL NETWORK MULTILINGUAL SPEECH RECOGNITION

発明審査中
1閲覧数
20請求項 · 3 独立
§ Ⅰ

案件概要

発明者

Purvi AGRAWAL; Vikas JOSHI; Basil ABRAHAM; Tejaswi SEERAM; Rupeshkumar Rasiklal MEHTA

IPC分類

G10L 15/16G10L 15/G10L 15/6G10L 15/22

CPC分類

G10L15/16G10L15/5G10L15/63G10L15/22

Systems, methods, and computer-readable storage devices are disclosed for improved recognition of multiple languages in audio data. One method including: receiving a trained split head multilingual neural network model, the trained split head multilingual neural network model including shared acoustic model layers and a plurality of projection layers, each projection layer of the plurality of projection layers corresponding to a language that the trained split head multilingual neural network model recognizes; receiving audio data, the audio data including speech in a plurality of languages in the audio data, the speech in the plurality of languages corresponding the language recognized by a projection layer of the plurality of projection layers of the trained split head multilingual neural network model; and classifying one or more languages of the speech of the audio data using the trained split head multilingual neural network model.

原文(中国語)

Systems, methods, and computer-readable storage devices are disclosed for improved recognition of multiple languages in audio data. One method including: receiving a trained split head multilingual neural network model, the trained split head multilingual neural network model including shared acoustic model layers and a plurality of projection layers, each projection layer of the plurality of projection layers corresponding to a language that the trained split head multilingual neural network model recognizes; receiving audio data, the audio data including speech in a plurality of languages in the audio data, the speech in the plurality of languages corresponding the language recognized by a projection layer of the plurality of projection layers of the trained split head multilingual neural network model; and classifying one or more languages of the speech of the audio data using the trained split head multilingual neural network model.

外部リソース