検索に戻る
案件記録

METHOD AND SYSTEM FOR PERFORMING VISION TASK USING PRE-TRAINED VISION-LANGUAGE TRANSFORMER

発明審査中
2閲覧数
20請求項 · 2 独立
§ Ⅰ

案件概要

発明者

Seung Hwan KIM; Jin Hyung KIM; Bum Soo KIM; Yeon Sik JO

IPC分類

G6N 3/96G6N 3/45G6T 11/60

CPC分類

G6N3/96G6N3/45G6T11/60

The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.

原文(中国語)

The present disclosure relates to a method and a system for promptly training a simplified vision-language transformer, in which large uncurated datasets are augmented (e.g., through image enlargement and/or masking, etc.) and vision-language transformers are pre-trained by reflecting, through a knowledge distillation framework, misaligned information between an augmented image and text upon the augmentation, thereby reducing both the necessary size of the utilized data set and data processing overhead.

外部リソース