CNIPA.AI
返回搜索
档案

PRESERVATION OF VISUAL CONTENT ACROSS MULTI-TURN DIALOGS WITH GENERATIVE MODEL(S)

发明专利审中
7浏览
20权利要求 · 3 独立
§ Ⅰ

卷宗概要

发明人

Ágoston Weisz; Alessandro Agostini; François-Xavier Aubet; Khalid Salama; Trevor Strohman; Ilia Akolzin; Petre Petrov

IPC 分类

G6F 40/284G6V 20/70

CPC 分类

G6F40/284G6V20/70

Implementations relate to handling visual content across a multi-turn dialog. A user input that includes natural language content and visual content is received during the dialog. If the visual content is being received for the first time in the dialog, the visual content is processed to generate a corresponding tokenized representation of the visual content. The corresponding tokenized representation can be cached in a database in association with the dialog, or in association with a user account of a user of the user query. If the visual content is subsequently referenced in the dialog, the corresponding tokenized representation of the visual content is retrieved from the database. The corresponding tokenized representation of the visual content, corresponding tokenized representations of natural language content, and optionally other metadata can be processed, using a generative model, to generate a response responsive to the user input.

原文(中文)

Implementations relate to handling visual content across a multi-turn dialog. A user input that includes natural language content and visual content is received during the dialog. If the visual content is being received for the first time in the dialog, the visual content is processed to generate a corresponding tokenized representation of the visual content. The corresponding tokenized representation can be cached in a database in association with the dialog, or in association with a user account of a user of the user query. If the visual content is subsequently referenced in the dialog, the corresponding tokenized representation of the visual content is retrieved from the database. The corresponding tokenized representation of the visual content, corresponding tokenized representations of natural language content, and optionally other metadata can be processed, using a generative model, to generate a response responsive to the user input.