Tangent rant. I'm skimming over some of the code at https://github.com/openai/gpt-2/blob/master/src/model.py and I can't help but feel frustrated at how unreadable this stuff is. 1. Why is it acceptable to have single-letter variable names everywhere? 2. There's little to almost no documentation in the code itself. It's unclear what the parameters of any given function mean. 3. There are magic constants everywhere. 4…
My professional observation (as ml researcher at big tech): These companies hire a lot of engineers straight out of undergrad/master's degrees. The interviews test leetcode knowledge, and today lots of degrees are heavy on Python-scripted ML homework. The result is companies with billion dollar funding and world-changing goals having a lot of their code look like complete spaghetti. And this is the engineers who are…
Many research scientist I've seen at large tech companies don't have good CS background and write code that's either unreadable, unscalable, unmaintainable, or buggy at times. They often have good ideas but don't do much software design.
Having been both at infrastructure teams and research teams, there are certain individuals in research orgs joining straight from school who think they are responsible with coming with a new and sexy thing and other engineers are responsible to run in production. It's like computing eigenvectors in Matlab or Python over a toy dataset and thinking you've done the bulk of the work of computing PageRank in production and should receive all the credit for a full search engine.
That attitude is a red flag to me.