検索に戻る
案件記録

SYSTEM AND METHOD OF CROSS-MODAL VISION-RADAR ALIGNMENT FOR OBJECT-LEVEL REPRESENTATION LEARNING

発明審査中
1閲覧数
20請求項 · 3 独立
§ Ⅰ

案件概要

発明者

Bingqing CHEN; Csaba DOMOKOS; Marcus PEREIRA; Kilian RAMBACH; João D. SEMEDO; Wan-Yi LIN; Leslie BERBERIAN

IPC分類

G6V 10/25

CPC分類

G6V10/25G6V2201/7

A method includes receiving a plurality of paired input images, wherein the paired images includes a first set of images from a first modality and a second set of images from a second modality, outputting a list of bounding boxes and labels in response to running an image-based object detection model, mapping each bounding box to a region of interest that is corresponding to the bounding box and associated with the second set of images, cropping the region of interest from the first and second set of images to generate a cropped first and second set of images, sending the cropped first set of images to a first encoder and a cropped second set of images to a second encoder, wherein the first encoder is configured for the first modality and the second encoder is configured for the second modality, outputting object-level embeddings for both the cropped first and second set of images utilizing encoders, identifying a loss function associated with the images, and in response to when a threshold is met, outputting final updated parameters.

原文(中国語)

A method includes receiving a plurality of paired input images, wherein the paired images includes a first set of images from a first modality and a second set of images from a second modality, outputting a list of bounding boxes and labels in response to running an image-based object detection model, mapping each bounding box to a region of interest that is corresponding to the bounding box and associated with the second set of images, cropping the region of interest from the first and second set of images to generate a cropped first and second set of images, sending the cropped first set of images to a first encoder and a cropped second set of images to a second encoder, wherein the first encoder is configured for the first modality and the second encoder is configured for the second modality, outputting object-level embeddings for both the cropped first and second set of images utilizing encoders, identifying a loss function associated with the images, and in response to when a threshold is met, outputting final updated parameters.

外部リソース