CNIPA.AI
返回搜索
档案

SOUNDING OBJECT FOCUSED SEGMENTATION FOR AN AUDIO-VISUAL SCENE

发明专利审中
1浏览
20权利要求 · 3 独立
§ Ⅰ

卷宗概要

发明人

Shi Yun Liang; Yuan Yuan Ding; Yong Sun; Zhong Yuan Sun; Jing Zhang

IPC 分类

G6V 10/80G6V 20/40H4R 5/4

CPC 分类

G6V10/806G6V20/46H4R5/4

An image feature vector for a given video frame is generated from a given video and an audio feature vector for audio of the given video is generated. A textual description of the given video frame is generated and textual feature vectors are generated from the textual description. A first set of audio features of the audio feature vector and visual features of the image feature vector are fused to generate fused audio-visual features. A second set of audio features of the audio feature vector and the textual feature vectors are fused to generate fused audio-text features. A final mask is generated based on the fused audio-visual features and the fused audio-text features.

原文(中文)

An image feature vector for a given video frame is generated from a given video and an audio feature vector for audio of the given video is generated. A textual description of the given video frame is generated and textual feature vectors are generated from the textual description. A first set of audio features of the audio feature vector and visual features of the image feature vector are fused to generate fused audio-visual features. A second set of audio features of the audio feature vector and the textual feature vectors are fused to generate fused audio-text features. A final mask is generated based on the fused audio-visual features and the fused audio-text features.