CNIPA.AI
返回搜索
档案

METHODS, APPARATUSES AND COMPUTER PROGRAM PRODUCTS FOR PROVIDING TUNING-FREE PERSONALIZED IMAGE GENERATION

发明专利审中
2浏览
20权利要求 · 3 独立
§ Ⅰ

卷宗概要

发明人

Zecheng He; Bo Sun; Juefei Xu; Animesh Sinha; Roshan Rajesh Sumbaly; Ning Zhang; Peizhao Zhang; Ankit Rajesh Ramchandani; Peter Vajda; Vincent Charles Cheung; Haoyu Ma

IPC 分类

G6T 11/60G6F 40/40G6V 10/26G6V 10/774G6V 40/16

CPC 分类

G6T11/60G6F40/40G6V10/26G6V10/774G6V40/161

A system and method to generate a target image from a reference image are provided. The system may receive, via a LDM, a reference image and a text prompt. The system may extract, via a trained vision encoder in the LDM, a vision control signal from an object in the reference image. The vision control signal indicates an identity of the object. The system may extract, via trained text encoders in the LDM, text control signals associated with the text prompt. The system may generate, via cross attention summation of an output of a vision cross attention unit(s) associated with the vision control signal and an output of text cross attention units associated with the text control signals, spatial features indicative of the reference image and the text prompt. The system may output, via a decoder in communication with the LDM, a target image based on the generated spatial features.

原文(中文)

A system and method to generate a target image from a reference image are provided. The system may receive, via a LDM, a reference image and a text prompt. The system may extract, via a trained vision encoder in the LDM, a vision control signal from an object in the reference image. The vision control signal indicates an identity of the object. The system may extract, via trained text encoders in the LDM, text control signals associated with the text prompt. The system may generate, via cross attention summation of an output of a vision cross attention unit(s) associated with the vision control signal and an output of text cross attention units associated with the text control signals, spatial features indicative of the reference image and the text prompt. The system may output, via a decoder in communication with the LDM, a target image based on the generated spatial features.