CNIPA.AI
검색으로 돌아가기
기록

Feedback Predictions for Machine-Learned Generative Models

발명심사 중
1조회수
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Junfeng He; Youwei Liang; Gang Li; Feng Yang; Junjie Ke; Peizhao Li; Vidhya Navalpakkam; Jiao Sun; Yang Li; Kai Jochen Kohlhoff; Jordi Pont-Tuset; Deepak Ramachandran

IPC 분류

G6T 11/60G6T 5/60G6T 5/77

CPC 분류

G6T11/60G6T5/60G6T5/77G6T2207/20081G6T2207/20084

Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models for feedback predictions for synthetic content. A machine-learned multimodal model is configured to generate a feature map based at least in part on fusion of image information and text information from a synthetic image and a text prompt. The model is configured to generate a set of text tokens based at least in part on fusion of the image information and the text information. The model is configured to generate at least one misalignment or implausibility heatmap based at least in part on the at least one feature map. The model is configured to generate at least one predicted misalignment sequence based at least in part on the set of text tokens.

원문 (중국어)

Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models for feedback predictions for synthetic content. A machine-learned multimodal model is configured to generate a feature map based at least in part on fusion of image information and text information from a synthetic image and a text prompt. The model is configured to generate a set of text tokens based at least in part on fusion of the image information and the text information. The model is configured to generate at least one misalignment or implausibility heatmap based at least in part on the at least one feature map. The model is configured to generate at least one predicted misalignment sequence based at least in part on the set of text tokens.