CNIPA.AI
검색으로 돌아가기
기록

METHOD AND SYSTEM FOR PERFORMING VISION TASK USING PRE-TRAINED VISION-LANGUAGE TRANSFORMER

발명심사 중
20청구항 · 2 독립항
§ Ⅰ

개요

발명자

Seung Hwan KIM; Jin Hyung KIM; Bum Soo KIM; Yeon Sik JO

IPC 분류

G6N 3/96G6N 3/45G6T 11/60

CPC 분류

G6N3/96G6N3/45G6T11/60

The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.

원문 (중국어)

The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.