TEXT-TO-SPEECH TRANSDUCER
卷宗概要
发明人
Vladimir Bataev; Subhankar Ghosh; Vitaly Lavrukhin; Boris Ginsburg
IPC 分类
CPC 分类
Disclosed are apparatuses, systems, and techniques that use a text-to-speech (TTS) transducer to perform TTS operations. The techniques include generating an initial input for a second model using an output of a first model. The techniques include generating, using the second model and the initial input, a first set of audio codes. The techniques include iteratively generating subsequent sets of audio codes using, at each iteration, the second model and a respective subsequent input for the second model. The respective subsequent input can reflect at least one previous set of audio codes generated by the second model.
原文(中文)
Disclosed are apparatuses, systems, and techniques that use a text-to-speech (TTS) transducer to perform TTS operations. The techniques include generating an initial input for a second model using an output of a first model. The techniques include generating, using the second model and the initial input, a first set of audio codes. The techniques include iteratively generating subsequent sets of audio codes using, at each iteration, the second model and a respective subsequent input for the second model. The respective subsequent input can reflect at least one previous set of audio codes generated by the second model.