CNIPA.AI
검색으로 돌아가기
기록

MULTI ATTENTION SPATIO-TEMPORAL MODEL FOR FINE-GRAINED VIDEO RECOGNITION

발명심사 중
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Marwa K. Qaraqe; Yin Yang; Elizabeth Varghese; Almiqdad Elzein

IPC 분류

G6V 10/82G6T 3/4007G6T 3/4046G6V 10/44G6V 10/77G6V 20/40

CPC 분류

G6V10/82G6T3/4007G6T3/4046G6V10/44G6V10/7715G6V20/46

A multi-attention spatio-temporal model for fine-grained video recognition is disclosed. This model offers a robust solution for fine-grained video recognition by addressing the intricate challenges of simultaneously considering complex spatial and temporal information, understanding temporal relationships between frames, dynamically allocating attention to informative spatial regions and temporal segments, and adapting to varying scales and resolutions. It empowers the model to not only pinpoint “where” and “when” to focus attention but also determine “how long” to make inferences, thereby enhancing overall performance. Experiments across diverse datasets demonstrate its efficiency in interpreting complex actions and scenes, enabling precise recognition. This innovation holds promise for a wide range of applications in computer vision facilitating more accurate and insightful video analysis.

원문 (중국어)

A multi-attention spatio-temporal model for fine-grained video recognition is disclosed. This model offers a robust solution for fine-grained video recognition by addressing the intricate challenges of simultaneously considering complex spatial and temporal information, understanding temporal relationships between frames, dynamically allocating attention to informative spatial regions and temporal segments, and adapting to varying scales and resolutions. It empowers the model to not only pinpoint “where” and “when” to focus attention but also determine “how long” to make inferences, thereby enhancing overall performance. Experiments across diverse datasets demonstrate its efficiency in interpreting complex actions and scenes, enabling precise recognition. This innovation holds promise for a wide range of applications in computer vision facilitating more accurate and insightful video analysis.