Charles Explorer logo
🇬🇧

Czech parliament meeting recordings as ASR training data

Publication at Faculty of Mathematics and Physics |
2020

Abstract

I present a way to leverage the stenographed recordings of the Czech parliament meetings for purposes of training a speech-to-text system. The article presents a method for scraping the data, acquiring word-level alignment and selecting reliable parts of the imprecise transcript.

Finally, I present an ASR system trained on these and other data.