PRESERVATION OF VISUAL CONTENT ACROSS MULTI-TURN DIALOGS WITH GENERATIVE MODEL(S)
Dossier Overview
Inventor
Ágoston Weisz; Alessandro Agostini; François-Xavier Aubet; Khalid Salama; Trevor Strohman; Ilia Akolzin; Petre Petrov
IPC Classification
CPC Classification
Implementations relate to handling visual content across a multi-turn dialog. A user input that includes natural language content and visual content is received during the dialog. If the visual content is being received for the first time in the dialog, the visual content is processed to generate a corresponding tokenized representation of the visual content. The corresponding tokenized representation can be cached in a database in association with the dialog, or in association with a user account of a user of the user query. If the visual content is subsequently referenced in the dialog, the corresponding tokenized representation of the visual content is retrieved from the database. The corresponding tokenized representation of the visual content, corresponding tokenized representations of natural language content, and optionally other metadata can be processed, using a generative model, to generate a response responsive to the user input.
Original (Chinese)
Implementations relate to handling visual content across a multi-turn dialog. A user input that includes natural language content and visual content is received during the dialog. If the visual content is being received for the first time in the dialog, the visual content is processed to generate a corresponding tokenized representation of the visual content. The corresponding tokenized representation can be cached in a database in association with the dialog, or in association with a user account of a user of the user query. If the visual content is subsequently referenced in the dialog, the corresponding tokenized representation of the visual content is retrieved from the database. The corresponding tokenized representation of the visual content, corresponding tokenized representations of natural language content, and optionally other metadata can be processed, using a generative model, to generate a response responsive to the user input.
External Resources