検索に戻る
案件記録

VISION-LANGUAGE MODEL FOR IMAGE CROPPING THROUGH IN-CONTEXT LEARNING

発明審査中
20請求項 · 2 独立
§ Ⅰ

案件概要

出願人

Google LLC

発明者

Feng Yang; Seunghyun Lee; Junjie Ke; Yinxiao Li; Junfeng He; Steven Hickson; Ekaterina Datsenko; Ming-Hsuan Yang; Irfan Essa

IPC分類

G6V 10/26G6F 16/532G6V 10/774G6V 10/776

CPC分類

G6V10/26G6F16/532G6V10/774G6V10/776

The technology provides for enhanced image cropping via in-context learning. It includes an efficient prompt retrieval mechanism for image cropping to automate the selection of in-context examples. It also includes an iterative refinement strategy to iteratively enhance the predicted crops. The image cropping framework is applicable to a wide range of cropping tasks, including free-form cropping, subject-aware cropping, and aspect ratio-aware cropping. The approach employs a trained large vision-language model associated with in-context learning. For instance, given an input image (whether from free-form, subject-aware or aspect ratio-aware cropping), the top-K semantically similar images from a dataset are retrieved as an in-context learning prompt. Then the in-context learning prompt is fed to a pretrained vision-language model to generate a set of crops. The crop candidates of the set are iteratively refined to yield a final output crop. The final output crop can then be applied to a downstream imaging task.

原文(中国語)

The technology provides for enhanced image cropping via in-context learning. It includes an efficient prompt retrieval mechanism for image cropping to automate the selection of in-context examples. It also includes an iterative refinement strategy to iteratively enhance the predicted crops. The image cropping framework is applicable to a wide range of cropping tasks, including free-form cropping, subject-aware cropping, and aspect ratio-aware cropping. The approach employs a trained large vision-language model associated with in-context learning. For instance, given an input image (whether from free-form, subject-aware or aspect ratio-aware cropping), the top-K semantically similar images from a dataset are retrieved as an in-context learning prompt. Then the in-context learning prompt is fed to a pretrained vision-language model to generate a set of crops. The crop candidates of the set are iteratively refined to yield a final output crop. The final output crop can then be applied to a downstream imaging task.

外部リソース