METHOD FOR GENERATING DIGITAL HUMAN VIDEO BASED ON LARGE MODEL, ELECTRONIC DEVICE, AND STORAGE MEDIUM
개요
출원인
Beijing Baidu Netcom Science Technology Co., Ltd.
발명자
Tian WU; Haifeng WANG; Hao TIAN; Wenquan WU; Dai DAI; Simei LIU; Li WANG; Hang ZHOU; Cong GAO; Qunyi XIE; Qingchang HAO
IPC 분류
CPC 분류
A method for generating a digital human video based on a large model, an electronic device, and a storage medium are provided, which relate to a field of artificial intelligence technologies, and may be applied to scenarios such as video livestreaming, advertisement production, and e-commerce sales. The method includes: acquiring a requirement information including an action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; processing the requirement information using a first large model to obtain a target script, where the target script includes a target speech segment text matching the action description information; and processing the target script and the action video segment using a second large model to obtain a target video for displaying a target digital human performing a speech delivery based on the target speech segment text while performing the specified action.
원문 (중국어)
A method for generating a digital human video based on a large model, an electronic device, and a storage medium are provided, which relate to a field of artificial intelligence technologies, and may be applied to scenarios such as video livestreaming, advertisement production, and e-commerce sales. The method includes: acquiring a requirement information including an action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; processing the requirement information using a first large model to obtain a target script, where the target script includes a target speech segment text matching the action description information; and processing the target script and the action video segment using a second large model to obtain a target video for displaying a target digital human performing a speech delivery based on the target speech segment text while performing the specified action.