検索に戻る
案件記録

MULTI ATTENTION SPATIO-TEMPORAL MODEL FOR FINE-GRAINED VIDEO RECOGNITION

発明審査中
4閲覧数
20請求項 · 3 独立
§ Ⅰ

案件概要

発明者

Marwa K. Qaraqe; Yin Yang; Elizabeth Varghese; Almiqdad Elzein

IPC分類

G6V 10/82G6T 3/4007G6T 3/4046G6V 10/44G6V 10/77G6V 20/40

CPC分類

G6V10/82G6T3/4007G6T3/4046G6V10/44G6V10/7715G6V20/46

A multi-attention spatio-temporal model for fine-grained video recognition is disclosed. This model offers a robust solution for fine-grained video recognition by addressing the intricate challenges of simultaneously considering complex spatial and temporal information, understanding temporal relationships between frames, dynamically allocating attention to informative spatial regions and temporal segments, and adapting to varying scales and resolutions. It empowers the model to not only pinpoint “where” and “when” to focus attention but also determine “how long” to make inferences, thereby enhancing overall performance. Experiments across diverse datasets demonstrate its efficiency in interpreting complex actions and scenes, enabling precise recognition. This innovation holds promise for a wide range of applications in computer vision facilitating more accurate and insightful video analysis.

原文(中国語)

A multi-attention spatio-temporal model for fine-grained video recognition is disclosed. This model offers a robust solution for fine-grained video recognition by addressing the intricate challenges of simultaneously considering complex spatial and temporal information, understanding temporal relationships between frames, dynamically allocating attention to informative spatial regions and temporal segments, and adapting to varying scales and resolutions. It empowers the model to not only pinpoint “where” and “when” to focus attention but also determine “how long” to make inferences, thereby enhancing overall performance. Experiments across diverse datasets demonstrate its efficiency in interpreting complex actions and scenes, enabling precise recognition. This innovation holds promise for a wide range of applications in computer vision facilitating more accurate and insightful video analysis.

外部リソース