CNIPA.AI
검색으로 돌아가기
기록

Multimodal Machine-Learned Models for Unified Attention and Response Predictions for Visual Content

발명심사 중
2조회수
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Junfeng He; Gang Li; Peizhao Li; Nachiappan Valliappan; Vidhya Navalpakkam; Yang Li; Kai Jochen Kohlhoff

IPC 분류

G6N 20/G6T 7/G6V 10/46

CPC 분류

G6N20/G6T7/2G6V10/462G6T2207/30168

Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models. A machine-learned multimodal model includes one or more embedding layers configured to generate one or more image tokens and one or more text tokens in response to the imagery and the text, a transformer encoder configured to receive the one or more image tokens and the one or more text tokens and generate one or more fused image tokens and one or more fused text tokens, a heatmap predictor configured to obtain the one or more fused image tokens and generate at least one image heatmap, and a sequence predictor configured to obtain the one or more fused image tokens and the one or more fused text tokens and generate a predicted sequence associated with the image.

원문 (중국어)

Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models. A machine-learned multimodal model includes one or more embedding layers configured to generate one or more image tokens and one or more text tokens in response to the imagery and the text, a transformer encoder configured to receive the one or more image tokens and the one or more text tokens and generate one or more fused image tokens and one or more fused text tokens, a heatmap predictor configured to obtain the one or more fused image tokens and generate at least one image heatmap, and a sequence predictor configured to obtain the one or more fused image tokens and the one or more fused text tokens and generate a predicted sequence associated with the image.