Multimodal Machine-Learned Models for Unified Attention and Response Predictions for Visual Content
案件概要
発明者
Junfeng He; Gang Li; Peizhao Li; Nachiappan Valliappan; Vidhya Navalpakkam; Yang Li; Kai Jochen Kohlhoff
IPC分類
CPC分類
Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models. A machine-learned multimodal model includes one or more embedding layers configured to generate one or more image tokens and one or more text tokens in response to the imagery and the text, a transformer encoder configured to receive the one or more image tokens and the one or more text tokens and generate one or more fused image tokens and one or more fused text tokens, a heatmap predictor configured to obtain the one or more fused image tokens and generate at least one image heatmap, and a sequence predictor configured to obtain the one or more fused image tokens and the one or more fused text tokens and generate a predicted sequence associated with the image.
原文(中国語)
Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models. A machine-learned multimodal model includes one or more embedding layers configured to generate one or more image tokens and one or more text tokens in response to the imagery and the text, a transformer encoder configured to receive the one or more image tokens and the one or more text tokens and generate one or more fused image tokens and one or more fused text tokens, a heatmap predictor configured to obtain the one or more fused image tokens and generate at least one image heatmap, and a sequence predictor configured to obtain the one or more fused image tokens and the one or more fused text tokens and generate a predicted sequence associated with the image.
外部リソース