Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

A study of auditory-filterbank based preprocessing for the generation of Warping-Invariant Features

Jan Rademacher, Alfred Mertins

Abstract

Auditory filterbanks have a long history in the preprocessing stage of automatic speech recognition systems, with the most prominent examples being the mel frequency cepstral coefficients (MFCCs). In this paper, we study the usefulness of auditory-filterbank analyses as a preprocessor for the generation of frequency-warping invariant features. The results indicate, that gammatone-filterbank analyses following the equivalent rectangular bandwidth (ERB) scale yield the most robust feature sets. The performance improvements are most significant when the vocal tract lengths in the training and test sets differ, which is important when, for example, children speech is to be recognized with a system that was mainly trained on adult data.
OriginalspracheEnglisch
Seiten1-5
Seitenumfang5
PublikationsstatusVeröffentlicht - 01.05.2006
VeranstaltungSpeech Recognition and Intrinsic Variation (SRIV2006)
- Toulouse, Frankreich
Dauer: 20.05.200620.05.2006

Tagung, Konferenz, Kongress

Tagung, Konferenz, KongressSpeech Recognition and Intrinsic Variation (SRIV2006)
Land/GebietFrankreich
OrtToulouse
Zeitraum20.05.0620.05.06

UN SDGs

Dieser Output leistet einen Beitrag zu folgendem(n) Ziel(en) für nachhaltige Entwicklung

  1. SDG 9 – Industrie, Innovation und Infrastruktur
    SDG 9 – Industrie, Innovation und Infrastruktur

Fingerprint

Untersuchen Sie die Forschungsthemen von „A study of auditory-filterbank based preprocessing for the generation of Warping-Invariant Features“. Zusammen bilden sie einen einzigartigen Fingerprint.

Zitieren