CNIPA.AI
검색으로 돌아가기
기록

VISION-LANGUAGE MODEL FOR IMAGE CROPPING THROUGH IN-CONTEXT LEARNING

발명심사 중
5조회수
20청구항 · 2 독립항
§ Ⅰ

개요

발명자

Feng Yang; Seunghyun Lee; Junjie Ke; Yinxiao Li; Junfeng He; Steven Hickson; Ekaterina Datsenko; Ming-Hsuan Yang; Irfan Essa

IPC 분류

G6V 10/26G6F 16/532G6V 10/774G6V 10/776

CPC 분류

G6V10/26G6F16/532G6V10/774G6V10/776

The technology provides for enhanced image cropping via in-context learning. It includes an efficient prompt retrieval mechanism for image cropping to automate the selection of in-context examples. It also includes an iterative refinement strategy to iteratively enhance the predicted crops. The image cropping framework is applicable to a wide range of cropping tasks, including free-form cropping, subject-aware cropping, and aspect ratio-aware cropping. The approach employs a trained large vision-language model associated with in-context learning. For instance, given an input image (whether from free-form, subject-aware or aspect ratio-aware cropping), the top-K semantically similar images from a dataset are retrieved as an in-context learning prompt. Then the in-context learning prompt is fed to a pretrained vision-language model to generate a set of crops. The crop candidates of the set are iteratively refined to yield a final output crop. The final output crop can then be applied to a downstream imaging task.

원문 (중국어)

The technology provides for enhanced image cropping via in-context learning. It includes an efficient prompt retrieval mechanism for image cropping to automate the selection of in-context examples. It also includes an iterative refinement strategy to iteratively enhance the predicted crops. The image cropping framework is applicable to a wide range of cropping tasks, including free-form cropping, subject-aware cropping, and aspect ratio-aware cropping. The approach employs a trained large vision-language model associated with in-context learning. For instance, given an input image (whether from free-form, subject-aware or aspect ratio-aware cropping), the top-K semantically similar images from a dataset are retrieved as an in-context learning prompt. Then the in-context learning prompt is fed to a pretrained vision-language model to generate a set of crops. The crop candidates of the set are iteratively refined to yield a final output crop. The final output crop can then be applied to a downstream imaging task.