METHOD AND SYSTEM FOR PERFORMING VISION TASK USING PRE-TRAINED VISION-LANGUAGE TRANSFORMER
개요
출원인
LG MANAGEMENT DEVELOPMENT INSTITUTE CO., LTD.
발명자
Seung Hwan KIM; Jin Hyung KIM; Bum Soo KIM; Yeon Sik JO
IPC 분류
CPC 분류
The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.
원문 (중국어)
The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.