Design and Evaluation of Automatic Speech Recognition Model Using Si Base Dl Model
DOI:
https://doi.org/10.71086/Keywords:
ASR, SI, DL, ML.Abstract
The signals of speech consist of sequences of sounds [1]. The transitions between these sounds serve as symbolic
representations of the information being conveyed. The way these sounds are arranged follows the rules of language.
Linguistics is the field that studies these rules and their application in human communication. Phonetics, on the other
hand, focuses on the sounds of speech. We can represent speech through the content of a given message. An acoustic
waveform is another way to characterize speech, capturing the signal that carries its message. For humans and their
environment, speech is vital for communication, which is why Automatic Speech Recognition (ASR) systems are
essential. There are various types of speech, including isolated words, continuous speech, and connected speech. The
study proposed for SER uses two supervised machine learning algorithms, support vector machines and decision trees,
enhanced with swarm intelligence optimization techniques, such as PSO and WOA, and a deep learning methodology
dubbed a deep belief network. The DBN automatically extracts features from the input speech signal after preprocessing.
Given that voice signals frequently lose a lot of information, making performance questionable, this
approach offers more dependable features.


