Earlier quoted context omitted.
This is a great analysis that really cleared up the impact of that memo for that me. You clearly know but for any readers who aren’t clear: that memo is the opinion of one engineer at google, and far from the (apparent) opinion of the relevant execs
It’s dead wrong, I suggest reading the original. _No one_ thinks models trained on more data would be about the same, since everything flows out of that premise…then throws Current Thing on top…it’s very unhelpful. The “patches” he refers to are LoRA and are treated as deus ex machina. They’re not, ex. playing with Stable Diffusion we can see they’re additive but they’re not nearly as good as training the original mo…
How are you defining "nearly as good", can you be more specific?
It's obvious full fine-tuning > [PEFT tuning] but to my understanding the gap isn't that significant, as reported in various papers. (specifically with respects to language models, I'm not familiar with diffusion models).