I mean, since GPT-4, I believe the RAM is no longer creating the miracle that the LLM performance scales directly with the model size. At least ChatGPT itself convinced me that any decent-sized company can create a GPT4 equivalent in terms of model size, but limited by service options, like memory cache and hallucination handling. Companies buy RAM simply to ride the stock hype. I am no expert, so this is a shallow t…
Re continuous fine-tuning: how do you avoid catastrophic forgetting in your proposal?
I put my ideas here in case you are interested:
https://github.com/SphericalCowww/ML_LunaLoRA