TECHNIQUES FOR ENHANCING SPEECH LANGUAGE MODELS USING DESCRIPTIVE SPEECH-TEXT ALIGNMENT
Dossier Overview
Applicant
NVIDIA CORPORATION
Inventor
Szu-Wei FU; Yu-Chiang WANG; Zhehuai CHEN; He HUANG; Boris GINSBURG
IPC Classification
CPC Classification
The disclosed method for generating a first depth map for responding to audio input includes processing the audio input using a trained encoder to generate a representation of the audio input, where the audio input includes speech; processing the representation of the audio input using a first trained adapter to generate one or more features; and processing the one or more features and text associated with the audio input using a trained language model to generate a response.
Original (Chinese)
The disclosed method for generating a first depth map for responding to audio input includes processing the audio input using a trained encoder to generate a representation of the audio input, where the audio input includes speech; processing the representation of the audio input using a first trained adapter to generate one or more features; and processing the one or more features and text associated with the audio input using a trained language model to generate a response.
External Resources