Have a personal or library account? Click to login

TEDxSK and JumpSK: A New Slovak Speech Recognition Dedicated Corpus

Open Access
|Jan 2018

Abstract

This paper describes a new Slovak speech recognition dedicated corpus built from TEDx talks and Jump Slovakia lectures. The proposed speech database consists of 220 talks and lectures in total duration of about 58 hours. Annotated speech database was generated automatically in an unsupervised manner by using acoustic speech segmentation based on principal component analysis and automatic speech transcription using two complementary speech recognition systems. The evaluation data consisting of 50 manually annotated talks and lectures in total duration of about 12 hours, has been created for evaluation of the quality of Slovak speech recognition. By unsupervised automatic annotation of TEDx talks and Jump Slovakia lectures we have obtained 21.26% of new speech segments with approximately 9.44% word error rate, suitable for retraining or adaptation of acoustic models trained beforehand.

DOI: https://doi.org/10.1515/jazcas-2017-0044 | Journal eISSN: 1338-4287 | Journal ISSN: 0021-5597
Language: English
Page range: 346 - 354
Published on: Jan 24, 2018
Published by: Slovak Academy of Sciences, Mathematical Institute
In partnership with: Paradigm Publishing Services
Publication frequency: 2 issues per year

© 2018 Ján Staš, Daniel Hládek, Peter Viszlay, Tomáš Koctúr, published by Slovak Academy of Sciences, Mathematical Institute
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.