Czech parliament meeting recordings as ASR training data

Publication at Faculty of Mathematics and Physics |

2020

Abstract

I present a way to leverage the stenographed recordings of the Czech parliament meetings for purposes of training a speech-to-text system. The article presents a method for scraping the data, acquiring word-level alignment and selecting reliable parts of the imprecise transcript.

Finally, I present an ASR system trained on these and other data.

Keywords

czech parliament meeting recordings training data

Czech parliament meeting recordings as ASR training data

Abstract

Keywords

Person