検索に戻る
案件記録

SOUNDING OBJECT FOCUSED SEGMENTATION FOR AN AUDIO-VISUAL SCENE

発明審査中
20請求項 · 3 独立
§ Ⅰ

案件概要

発明者

Shi Yun Liang; Yuan Yuan Ding; Yong Sun; Zhong Yuan Sun; Jing Zhang

IPC分類

G6V 10/80G6V 20/40H4R 5/4

CPC分類

G6V10/806G6V20/46H4R5/4

An image feature vector for a given video frame is generated from a given video and an audio feature vector for audio of the given video is generated. A textual description of the given video frame is generated and textual feature vectors are generated from the textual description. A first set of audio features of the audio feature vector and visual features of the image feature vector are fused to generate fused audio-visual features. A second set of audio features of the audio feature vector and the textual feature vectors are fused to generate fused audio-text features. A final mask is generated based on the fused audio-visual features and the fused audio-text features.

原文(中国語)

An image feature vector for a given video frame is generated from a given video and an audio feature vector for audio of the given video is generated. A textual description of the given video frame is generated and textual feature vectors are generated from the textual description. A first set of audio features of the audio feature vector and visual features of the image feature vector are fused to generate fused audio-visual features. A second set of audio features of the audio feature vector and the textual feature vectors are fused to generate fused audio-text features. A final mask is generated based on the fused audio-visual features and the fused audio-text features.

外部リソース