検索に戻る
案件記録

Multimodal Machine-Learned Models for Unified Attention and Response Predictions for Visual Content

発明審査中
1閲覧数
20請求項 · 3 独立
§ Ⅰ

案件概要

発明者

Junfeng He; Gang Li; Peizhao Li; Nachiappan Valliappan; Vidhya Navalpakkam; Yang Li; Kai Jochen Kohlhoff

IPC分類

G6N 20/G6T 7/G6V 10/46

CPC分類

G6N20/G6T7/2G6V10/462G6T2207/30168

Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models. A machine-learned multimodal model includes one or more embedding layers configured to generate one or more image tokens and one or more text tokens in response to the imagery and the text, a transformer encoder configured to receive the one or more image tokens and the one or more text tokens and generate one or more fused image tokens and one or more fused text tokens, a heatmap predictor configured to obtain the one or more fused image tokens and generate at least one image heatmap, and a sequence predictor configured to obtain the one or more fused image tokens and the one or more fused text tokens and generate a predicted sequence associated with the image.

原文(中国語)

Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models. A machine-learned multimodal model includes one or more embedding layers configured to generate one or more image tokens and one or more text tokens in response to the imagery and the text, a transformer encoder configured to receive the one or more image tokens and the one or more text tokens and generate one or more fused image tokens and one or more fused text tokens, a heatmap predictor configured to obtain the one or more fused image tokens and generate at least one image heatmap, and a sequence predictor configured to obtain the one or more fused image tokens and the one or more fused text tokens and generate a predicted sequence associated with the image.

外部リソース