Earlier quoted context omitted.
Like most machine-learning applications, the source code isn't the interesting part, the data is. Google started by training on millions of phrases from Google 411, and then they've been able to continue training anytime someone issues a voice command to an Android device. They have orders of magnitude more data than you could fit into a GitHub repository.
Couldn't you just download subtitles for old movies and train using those?
https://asciinema.org/a/93ihv5er83mrpihh9i5m3s9o5
Congress has video for all of its sessions and it is transcribed. So does the Supreme Court (though not timestamped).