Earlier quoted context omitted.
The point is, OpenAI spent a lot of money on training on all these copyrighted materials ordinary individuals/companies don't have access to, so replicating their effort would mean that you either 1) spend a ridiculous amount of money, 2) use Library Genesis (and still pay millions for GPU usage). So we have very little choice now. Open Source LLMs might be getting close to ChatGPT3 (opinions vary), but OpenAI is sti…
the choice is to live 2 years behind (e.g. integrate the open source stuff and ride that wave of improvement). for businesses in a competitive space, that’s perhaps untenable. but for individuals and anywhere else where this stuff is just a “nice to have”, that’s really just the long-term sustainable approach. it reminds me of a choice like “do i host my website on a Windows Server, or a Linux box” at a time when bot…
That's one world - there is another where the time gap grows a lot more as the compute and training requirements continue to rise.
Microsoft will probably be willing to spend multiple billions in compute to help train GPT5, so it depends how much investment open source projects can get to compete. Seems like it's down to Meta, but it depends if they can continue to justify releasing future models as Open Source considering the investment required, or what licensing looks like.