Unknown

Dataset Information

Data augmentation and deep neural networks for the classification of Pakistani racial speakers recognition.


ABSTRACT: Speech emotion recognition (SER) systems have evolved into an important method for recognizing a person in several applications, including e-commerce, everyday interactions, law enforcement, and forensics. The SER system's efficiency depends on the length of the audio samples used for testing and training. However, the different suggested models successfully obtained relatively high accuracy in this study. Moreover, the degree of SER efficiency is not yet optimum due to the limited database, resulting in overfitting and skewing samples. Therefore, the proposed approach presents a data augmentation method that shifts the pitch, uses multiple window sizes, stretches the time, and adds white noise to the original audio. In addition, a deep model is further evaluated to generate a new paradigm

SUBMITTER: Amjad A 

PROVIDER: S-EPMC9454772 | biostudies-literature | 2022

REPOSITORIES: biostudies-literature

altmetric image

Publications

Sorry, this publication's infomation has not been loaded in the Indexer, please go directly to PUBMED or Altmetric.

Similar Datasets