CNIPA.AI
검색으로 돌아가기
기록

POSITIONAL ENCODINGS FOR PERCEPTION FUNCTIONS IN AUTOMATED DRIVING SYSTEMS

발명심사 중
1조회수
14청구항 · 2 독립항
§ Ⅰ

개요

발명자

Willem VERBEKE; Joakim JOHNANDER

IPC 분류

G6T 7/73B60W 50/

CPC 분류

G6T7/74B60W50/97B60W2420/403G6T2207/20081G6T2207/20084G6T2207/30252

A method for making perception predictions for a perception functionality in an automated driving system of a vehicle is disclosed. The method includes generating 2D position information of an image captured by a vehicle-mounted camera. The 2D position information indicates a position of each pixel out of a plurality of pixels of the image, or a position of each patch out of a plurality of patches of the image in the 2D reference frame of the image. Then, feeding the generated 2D position information, extrinsic parameters of the vehicle-mounted camera, intrinsic parameters of the vehicle-mounted camera, and distortion parameters of the vehicle-mounted camera to a multilayer perceptron which process the feed data and output 3D positional encodings. The method further includes feeding the image data and the 3D positional encodings to a transformer network for generating a prediction output in a 3D/2D reference frame of the vehicle.

원문 (중국어)

A method for making perception predictions for a perception functionality in an automated driving system of a vehicle is disclosed. The method includes generating 2D position information of an image captured by a vehicle-mounted camera. The 2D position information indicates a position of each pixel out of a plurality of pixels of the image, or a position of each patch out of a plurality of patches of the image in the 2D reference frame of the image. Then, feeding the generated 2D position information, extrinsic parameters of the vehicle-mounted camera, intrinsic parameters of the vehicle-mounted camera, and distortion parameters of the vehicle-mounted camera to a multilayer perceptron which process the feed data and output 3D positional encodings. The method further includes feeding the image data and the 3D positional encodings to a transformer network for generating a prediction output in a 3D/2D reference frame of the vehicle.