CNIPA.AI
返回搜索
档案

METHOD AND SYSTEM FOR PERFORMING VISION TASK USING PRE-TRAINED VISION-LANGUAGE TRANSFORMER

发明专利审中
1浏览
20权利要求 · 2 独立
§ Ⅰ

卷宗概要

发明人

Seung Hwan KIM; Jin Hyung KIM; Bum Soo KIM; Yeon Sik JO

IPC 分类

G6N 3/96G6N 3/45G6T 11/60

CPC 分类

G6N3/96G6N3/45G6T11/60

The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.

原文(中文)

The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.