Feedback Predictions for Machine-Learned Generative Models
卷宗概要
发明人
Junfeng He; Youwei Liang; Gang Li; Feng Yang; Junjie Ke; Peizhao Li; Vidhya Navalpakkam; Jiao Sun; Yang Li; Kai Jochen Kohlhoff; Jordi Pont-Tuset; Deepak Ramachandran
IPC 分类
CPC 分类
Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models for feedback predictions for synthetic content. A machine-learned multimodal model is configured to generate a feature map based at least in part on fusion of image information and text information from a synthetic image and a text prompt. The model is configured to generate a set of text tokens based at least in part on fusion of the image information and the text information. The model is configured to generate at least one misalignment or implausibility heatmap based at least in part on the at least one feature map. The model is configured to generate at least one predicted misalignment sequence based at least in part on the set of text tokens.
原文(中文)
Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models for feedback predictions for synthetic content. A machine-learned multimodal model is configured to generate a feature map based at least in part on fusion of image information and text information from a synthetic image and a text prompt. The model is configured to generate a set of text tokens based at least in part on fusion of the image information and the text information. The model is configured to generate at least one misalignment or implausibility heatmap based at least in part on the at least one feature map. The model is configured to generate at least one predicted misalignment sequence based at least in part on the set of text tokens.