Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

How to Efficiently Use Color and Temporal Information for Video Understanding

Yaxin Hu*, Erhardt Barth

*Korrespondierende/r Autor/-in für diese Arbeit

Abstract

The modeling of temporal dependencies, and the associated computational load, remain challenges in video understanding. We here focus on using a more efficient sampling of color and temporal information. We sample color not from the same frame but from different consecutive frames to capture richer temporal information without increasing the computational load. We demonstrate the effectiveness of our approach for 2D-CNNs, 3D-CNNs, and Transformers, for which we obtain significant performance improvements on two benchmarks. The improvements are 2.43% on UCF101 and 4.55% on HMDB51 for the ResNet18, 10.28% and 7.12% for the 3D-ResNet18, and 15.11% and 13.71% for the UniFormerV2. These improvements are obtained without additional costs by just changing the way color is sampled.
OriginalspracheEnglisch
TitelLecture Notes in Computer Science
Seitenumfang14
Band15293
Herausgeber (Verlag)Springer Nature Singapore
Erscheinungsdatum02.12.2024
Seiten413-426
ISBN (Print)978-981-96-6598-3
ISBN (elektronisch)978-981-96-6596-9
PublikationsstatusVeröffentlicht - 02.12.2024

UN SDGs

Dieser Output leistet einen Beitrag zu folgendem(n) Ziel(en) für nachhaltige Entwicklung

  1. SDG 3 – Gesundheit und Wohlergehen
    SDG 3 – Gesundheit und Wohlergehen
  2. SDG 9 – Industrie, Innovation und Infrastruktur
    SDG 9 – Industrie, Innovation und Infrastruktur

Fingerprint

Untersuchen Sie die Forschungsthemen von „How to Efficiently Use Color and Temporal Information for Video Understanding“. Zusammen bilden sie einen einzigartigen Fingerprint.

Zitieren