Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Video Understanding Using 2D-CNNs on Salient Spatio-Temporal Slices

Yaxin Hu*, Erhardt Barth

*Korrespondierende/r Autor/-in für diese Arbeit

Abstract

Video understanding remains a challenge even with advanced deep-learning methods that typically sample a few frames from which spatial and temporal features are extracted. Such down-sampling often leads to the loss of critical temporal information. Moreover, current state-of-the-art methods involve high computational costs. 2D Convolutional Neural Networks (2D-CNNs) have proven to be effective at capturing spatial features of images, but cannot make use of temporal information. To address these challenges, we propose to use 2D-CNNs not only on images, i.e. xy-slices of the video, but on salient spatio-temporal xt and yt slices to efficiently capture both spatial and temporal information of the entire video. As 2D-CNNs are known to extract local spatial orientation in xy, they can now extract motion, which is a local orientation in xt and yt. We complement the approach with a simple strategy for sampling the most informative slices and show that we can outperform alternative approaches in a number of tasks, especially in cases in which the actions are defined by their dynamics, i.e., by spatio-temporal patterns.
OriginalspracheEnglisch
TitelLecture Notes in Computer Science : Artificial Neural Networks and Machine Learning – ICANN 2024
Band15018
Herausgeber (Verlag)Springer, Cham
Erscheinungsdatum17.09.2024
Seiten256-270
ISBN (Print)978-3-031-72337-7
ISBN (elektronisch)978-3-031-72338-4
PublikationsstatusVeröffentlicht - 17.09.2024

UN SDGs

Dieser Output leistet einen Beitrag zu folgendem(n) Ziel(en) für nachhaltige Entwicklung

  1. SDG 3 – Gesundheit und Wohlergehen
    SDG 3 – Gesundheit und Wohlergehen
  2. SDG 4 – Qualitativ hochwertige Bildung
    SDG 4 – Qualitativ hochwertige Bildung
  3. SDG 9 – Industrie, Innovation und Infrastruktur
    SDG 9 – Industrie, Innovation und Infrastruktur
  4. SDG 11 – Nachhaltige Städte und Gemeinschaften
    SDG 11 – Nachhaltige Städte und Gemeinschaften
  5. SDG 12 – Verantwortungsvoller Konsum und Produktion
    SDG 12 – Verantwortungsvoller Konsum und Produktion
  6. SDG 14 – Lebensraum Wasser
    SDG 14 – Lebensraum Wasser
  7. SDG 15 – Lebensraum Land
    SDG 15 – Lebensraum Land

Fingerprint

Untersuchen Sie die Forschungsthemen von „Video Understanding Using 2D-CNNs on Salient Spatio-Temporal Slices“. Zusammen bilden sie einen einzigartigen Fingerprint.

Zitieren