CNIPA.AI
返回搜索
档案

MULTI ATTENTION SPATIO-TEMPORAL MODEL FOR FINE-GRAINED VIDEO RECOGNITION

发明专利审中
2浏览
20权利要求 · 3 独立
§ Ⅰ

卷宗概要

发明人

Marwa K. Qaraqe; Yin Yang; Elizabeth Varghese; Almiqdad Elzein

IPC 分类

G6V 10/82G6T 3/4007G6T 3/4046G6V 10/44G6V 10/77G6V 20/40

CPC 分类

G6V10/82G6T3/4007G6T3/4046G6V10/44G6V10/7715G6V20/46

A multi-attention spatio-temporal model for fine-grained video recognition is disclosed. This model offers a robust solution for fine-grained video recognition by addressing the intricate challenges of simultaneously considering complex spatial and temporal information, understanding temporal relationships between frames, dynamically allocating attention to informative spatial regions and temporal segments, and adapting to varying scales and resolutions. It empowers the model to not only pinpoint “where” and “when” to focus attention but also determine “how long” to make inferences, thereby enhancing overall performance. Experiments across diverse datasets demonstrate its efficiency in interpreting complex actions and scenes, enabling precise recognition. This innovation holds promise for a wide range of applications in computer vision facilitating more accurate and insightful video analysis.

原文(中文)

A multi-attention spatio-temporal model for fine-grained video recognition is disclosed. This model offers a robust solution for fine-grained video recognition by addressing the intricate challenges of simultaneously considering complex spatial and temporal information, understanding temporal relationships between frames, dynamically allocating attention to informative spatial regions and temporal segments, and adapting to varying scales and resolutions. It empowers the model to not only pinpoint “where” and “when” to focus attention but also determine “how long” to make inferences, thereby enhancing overall performance. Experiments across diverse datasets demonstrate its efficiency in interpreting complex actions and scenes, enabling precise recognition. This innovation holds promise for a wide range of applications in computer vision facilitating more accurate and insightful video analysis.