Skip to main content
Have a personal or library account? Click to login
Applicability of End-to-End Deep Neural Architecture to Sinhala Speech Recognition Cover

Applicability of End-to-End Deep Neural Architecture to Sinhala Speech Recognition

Open Access
|May 2024

Abstract

This research presents a study on the application of end-to-end deep learning models for Automatic Speech Recognition in the Sinhala language, which is characterized by its high inflection and limited resources.We explore two e2e architectures, namely the e2e Lattice-Free Maximum Mutual Information model and the Recurrent Neural Network model, using a restricted dataset. Statistical models with 40 hours of training data are established as baselines for evaluation. Our pretrained endto-end Automatic Speech Recognition models achieved a Word Error Rate of 23.38% by far the best word-error-rate achieved for low resourced Sinhala Language. Our models demonstrate greater contextual independence and faster processing, making them more suitable for general-purpose speech-to-text translation in Sinhala.
Language: English
Page range: 17 - 21
Published on: May 31, 2024
Published by: University of Colombo School of Computing
In partnership with: Paradigm Publishing Services

© 2024 Buddhi Gamage, Randil Pushpananda, Thilini Nadungodage, Ruvan Weerasinghe, published by University of Colombo School of Computing
This work is licensed under the Creative Commons Attribution 4.0 License.