CNIPA.AI
검색으로 돌아가기
기록

SOUNDING OBJECT FOCUSED SEGMENTATION FOR AN AUDIO-VISUAL SCENE

발명심사 중
3조회수
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Shi Yun Liang; Yuan Yuan Ding; Yong Sun; Zhong Yuan Sun; Jing Zhang

IPC 분류

G6V 10/80G6V 20/40H4R 5/4

CPC 분류

G6V10/806G6V20/46H4R5/4

An image feature vector for a given video frame is generated from a given video and an audio feature vector for audio of the given video is generated. A textual description of the given video frame is generated and textual feature vectors are generated from the textual description. A first set of audio features of the audio feature vector and visual features of the image feature vector are fused to generate fused audio-visual features. A second set of audio features of the audio feature vector and the textual feature vectors are fused to generate fused audio-text features. A final mask is generated based on the fused audio-visual features and the fused audio-text features.

원문 (중국어)

An image feature vector for a given video frame is generated from a given video and an audio feature vector for audio of the given video is generated. A textual description of the given video frame is generated and textual feature vectors are generated from the textual description. A first set of audio features of the audio feature vector and visual features of the image feature vector are fused to generate fused audio-visual features. A second set of audio features of the audio feature vector and the textual feature vectors are fused to generate fused audio-text features. A final mask is generated based on the fused audio-visual features and the fused audio-text features.