METHOD AND SYSTEM FOR PERFORMING VISION TASK USING PRE-TRAINED VISION-LANGUAGE TRANSFORMER
卷宗概要
申请人
LG MANAGEMENT DEVELOPMENT INSTITUTE CO., LTD.
发明人
Seung Hwan KIM; Jin Hyung KIM; Bum Soo KIM; Yeon Sik JO
IPC 分类
CPC 分类
The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.
原文(中文)
The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.