Live data from Hacker News

Training LLMs from ground zero as a startup

yitay.net

71–80 of 125 posts

Re: Training LLMs from ground zero as a startup

#72
post #55
post #26

Earlier quoted context omitted.

Great too see you on here. Love Latent Space podcast.

aw thank you for listening. some weeks its very much a labor of love lol. no events planned near term but come to the big shindig in june https://ti.to/software-3/ai-engineer-worlds-fair . last year's summit was the first time i really understood how much of a reach we have and how many good AI people we've managed to gather as friends.

I love it as well, it’s a fantastic resource :)

Re: Training LLMs from ground zero as a startup

#73

I’m wondering if the title should read “from the ground up” instead of “ground zero”? https://en.wikipedia.org/wiki/Hypocenter

https://www.merriam-webster.com/dictionary/ground%20zero It is a perfectly acceptable use of the idiom.

Acceptable, but maybe not perfectly.

Re: Training LLMs from ground zero as a startup

#76
post #2

for context Yi Tay was Tech Lead on Google PaLM, UL2, Flan, Bard, etc and now is cofoudner at Reka (which has shipped some v interesting small multimodal models that have featured on here). I prompted him for this post as an ex-Googler now training LLMs as an independent startup https://twitter.com/YiTayML/status/1765105066263052718 our conversation was recorded here https://sub.thursdai.news/p/thursdai-feb-15-2024-o…

Is he the person after the Yi LLM model?

Re: Training LLMs from ground zero as a startup

#77

> To be very frank, I would have to say the quality of codebases externally significantly lag behind those I’ve been used to at Google Haven't worked at Google, anyone else share this sentiment? I always feel like working with Google code is typically not idiomatic and super difficult to go "under the hood" if anything isn't precisely on the happy path.

I thought the quality was pretty high, largely because there were a lot of rails constraining how code should be written. Most of the code I dealt with was written using somewhat rigid (but generally well-designed) frameworks with programmatically-enforced style guides. Also, most work seemed to involve some balance of junior and more experienced people, which helped keep quality higher. Outside of Google, I've seen…

The thing that impressed me most about Google was the encoding-of-cultural-norms-in-various-CI-jobs.

It lets them extract usable SWE horsepower from pretty much anyone who steps inside and at least tries to be useful and not just coast. They can ingest a startup engineer, someone who's been a mid-tier enterprise codemonkey, yr mythical 10xer, the whole statistical gamut.

Re: Training LLMs from ground zero as a startup

#79
> In the end it took us only a very small number of smaller scale & shorter ablation runs to get to the strong 21B Reka Flash and 7B edge model (and also our upcoming largest core model). Finding a solid recipe with a very limited number of runs is challenging and requires changing many variables at once given the ridiculously enormous search space. In order to do this, one has to abandon the systematicity of Bigtech and rely a lot on “Yolo”, gut feeling and instinct.

> Thankfully, I (and many of us in the team) have built up this intuition quite a bit in our ML careers to get it right within a substantially short amount of tries. While we’ve trained really good models before in our previous jobs, differences in training infrastructure, data, incorporation of new ideas and other environmental issues can still cause non-trivial differences in outcomes. That said, a strong prior helps to significantly cut down the search space and is probably one of the easiest explanations to why we were able to train really strong models with so few trials, resources and experimentation.

Post reply on HN