CNIPA.AI
검색으로 돌아가기
기록

METHOD, ELECTRONIC DEVICE, AND COMPUTER PROGRAM PRODUCT FOR GENERATING VIDEO

발명심사 중
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Zijia Wang; Zhisong Liu; Jiacheng Ni; Zhen Jia

IPC 분류

G6T 13/20G6T 3/40G6T 5/70G6T 13/40G6T 13/80G10L 15/2G10L 15/6G10L 15/18G10L 25/57

CPC 분류

G6T13/205G6T3/40G6T5/70G6T13/40G6T13/80G10L15/2G10L15/63G10L15/1815G10L25/57G6T2207/10016G6T2207/20081

A method includes obtaining a reference image and a reference speech, the reference image specifying a head of a target object in the video, and the reference speech specifying a voice of the target object; and generating, based on the reference image and the reference speech, a fusion vector by combining a feature of the head and a feature of the voice. The method further includes generating, based on the fusion vector, a plurality of video frames in a video that represents the target object speaking in a timbre of the reference speech by denoising a plurality of initial frames including noise; and generating the video based on the plurality of video frames. In embodiments of the present disclosure, a video in which a semantic feature and a speaking style of the target object are merged can be generated, and the resolution and quality of the generated video are enhanced.

원문 (중국어)

A method includes obtaining a reference image and a reference speech, the reference image specifying a head of a target object in the video, and the reference speech specifying a voice of the target object; and generating, based on the reference image and the reference speech, a fusion vector by combining a feature of the head and a feature of the voice. The method further includes generating, based on the fusion vector, a plurality of video frames in a video that represents the target object speaking in a timbre of the reference speech by denoising a plurality of initial frames including noise; and generating the video based on the plurality of video frames. In embodiments of the present disclosure, a video in which a semantic feature and a speaking style of the target object are merged can be generated, and the resolution and quality of the generated video are enhanced.