CNIPA.AI
검색으로 돌아가기
기록

TECHNIQUES FOR ENHANCING SPEECH LANGUAGE MODELS USING DESCRIPTIVE SPEECH-TEXT ALIGNMENT

발명심사 중
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Szu-Wei FU; Yu-Chiang WANG; Zhehuai CHEN; He HUANG; Boris GINSBURG

IPC 분류

G10L 15/183G10L 15/2G10L 15/22

CPC 분류

G10L15/183G10L15/2G10L15/22

The disclosed method for generating a first depth map for responding to audio input includes processing the audio input using a trained encoder to generate a representation of the audio input, where the audio input includes speech; processing the representation of the audio input using a first trained adapter to generate one or more features; and processing the one or more features and text associated with the audio input using a trained language model to generate a response.

원문 (중국어)

The disclosed method for generating a first depth map for responding to audio input includes processing the audio input using a trained encoder to generate a representation of the audio input, where the audio input includes speech; processing the representation of the audio input using a first trained adapter to generate one or more features; and processing the one or more features and text associated with the audio input using a trained language model to generate a response.