An Intelligent Deep Learning Framework For Spoken Language Identification

Authors: K. Rajkumar, A. Swarshitha

Abstract: Modern speech processing applications rely heavily on spoken language identification. This includes intelligent translation systems, voice-based human-computer interaction, and multilingual speech recognition. Communication technologies are enhanced by accurate voice recognition, which allows for efficient processing of multilingual speech data. The five most common Indian languages—Hindi, Bengali, Tamil, English, and Gujarati—are identified using a framework based on deep learning in this study. To classify languages, the suggested technique feeds a Deep Neural Network (DNN) with audio data extracted from speech signals using Mel-Frequency Cepstral Coefficients (MFCCs). To enhance classification performance, audio samples are preprocessed, features are extracted, and models are trained using publically accessible speech datasets. Common performance measures, including as F1-score, recall, accuracy, and precision, are used to assess the efficacy of the suggested model. The results of the experiments show that the deep learning model successfully differentiates between the target languages and shows impressive generalisation when applied to new audio data. In addition, the published web application is built using Django, which incorporates the trained model. Users may access the results of language recognition in real-time by uploading voice recordings and interacting with the interface. With its extensible design, the suggested framework may accommodate more languages and practical speech processing applications while simultaneously providing an efficient, scalable, and workable solution for multilingual spoken language recognition.

DOI: http://doi.org/10.5281/zenodo.21374453

Leave a Reply

Your email address will not be published. Required fields are marked *