Earlier quoted context omitted.
Selling 20 million Quest 2 headsets is a pretty good outcome for the so-called Metaverse fiasco.
I think they spent 36 billion on metaverse stuff at last count, so no, not really.
The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
281–290 of 527 posts
Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
#282Earlier quoted context omitted.
I tried that with RTX 4090 as the primary card and 3090 as eGPU over Thunderbolt. It works, but the inference is very slow, presumably because it has to pump all that data back and forth between the two (and Thunderbolt isn't fast enough to keep up even with 3090 by itself in games). In fact, even running 30B across two GPUs in 8-bit mode like that was slower than running it on one GPU in 4-bit. My takeaway is that i…
Looking at the top end H100 80GB systems with NVLink from HPC vendors and it occurred to me we are about to swing back to massive almost mainframe like form-factor systems, giant bus, like old expandable qbus in the 80s but this time for GPUs. What I mean is they have systems with 8x cards but given the compute requirements of these huge LLMs probably systems with 32+ all on a dedicated memory bus (NVLink) are what w…
Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
#283Just some examples of things you can do.
How about create a D&D (or any RPG) game, the NPC can be the dungeon master, creating monsters/loot and rendering pics in real time. Add additional NPC characters to join your party and actually take turns, you could play solo adventures. Even play via microphone, if the voice extension gets modified, you could have each character have its own voice via tts.
The extensions are opensource to make anything you want, connect it to web or any service. Have the npc chat avatars trigger on actions.
You can even train models, want to create a NPC based on a book? Feed it book series, tweak the personality, and you can chat with them, or make up new stories. The training model interface is included.
Or for adults you could even create a virtual partner, or any type of NPC/avatar you want. Have them text you stable diffusion pics, chat with you on sms. etc.
AND, the thing is, its out NOW on github with text-generation-webui. I was able to create a D&D dungeon master with stable diffusion in about 10 minutes. I did already have stable diffusion running thou, just enabled the api.
I can't wait to see how this new amazing software can take off to form new ideas, games, technology.
Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
#284Earlier quoted context omitted.
Is it possible to build systems with multiple GPUs to run the 65B or larger when they appear? I’m not really sure and looking for clarification from anyone who knows. My understanding is it is possible to split the layers between the GPUs so a system with 4 high end consumer GPUs might work well.
I tried that with RTX 4090 as the primary card and 3090 as eGPU over Thunderbolt. It works, but the inference is very slow, presumably because it has to pump all that data back and forth between the two (and Thunderbolt isn't fast enough to keep up even with 3090 by itself in games). In fact, even running 30B across two GPUs in 8-bit mode like that was slower than running it on one GPU in 4-bit. My takeaway is that i…
Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
#285I've spent an embarassing amount of time since the llamas leaked playing with them, the tools to run them, and writing wrappers for them. They are technically alternatives in the sense that they're incomparably better chat bots than anything in the past. But at least for the 30B and under versions (65B is too big for me to run), no matter what fine tuning is done (alpaca, gpt4all, vicuna, etc), the llamas themselves…
I haven’t seen to much discussion of what’s possible at various sizes for an early stage start up, which is a discussion I’d expect to see on yc. Clearly a company with $5-5MM in the bank can’t train a competitive LLM from scratch but what would it cost to fine tune and/or run a 65B parameter model or a hypothetical future open source 165B parameter model?
Wait, are we sure?
I'm going to make the massive mistake of assuming we're compute bound instead of memory bound, and assume we can train at FP16 (which is a bad assumption because, of course, you're doing calculus where the little pieces you're adding up could get rounded to zero at FP16 pretty easily... although mixed precision FP32/FP16 training is possible).
Consumer GPUs like GeForce RTX 4090 can do 3e14 flop/s under certain conditions with fp16. They retailed for about $1600. It took reportedly 3e23 flop to train GPT-3. A year is 3e7 seconds. So the upfront cost of retail GPUs doing 3e23 fp16 operations in a single year is potentially as low as ~$50k (and about $20k worth of electricity). (FP32 peak is about a factor of 4 worse, so ~$200k.)
So it's not actually impossible to imagine a particularly clever approach to training that could maybe achieve competitive LLM training for less than $5 million in hardware costs. (except for the fact that compute isn't really the bottleneck, memory really is.)
Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
#286Earlier quoted context omitted.
This is an interesting argument as it's easy to apply it nearly universally to any example of learning. What sort of evidence would convince you that it is learning?
Since when training and fine-tuning isn't learning ? Individual sessions of LLMs are not learning, but models as products surely are - the feedback loop is just iterated manually.
To me the LLM loophole/"hack" closings just feel like a human vs human cat&mouse game with some Chat UI in the middle.
Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
#287Earlier quoted context omitted.
The evolution of answers from version to version makes it clear there are insane amounts of manual fine tunings happening. I think this is largely overlooked by the "its learning" crowd.
This is an interesting argument as it's easy to apply it nearly universally to any example of learning. What sort of evidence would convince you that it is learning?
Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
#288I'm a bit worried the LLaMA leak will make the labs much more cautious about who they distribute models to for future projects, closing down things even more. I've had tons of fun implementing LLaMA, learning and playing around with variations like Vicuna. I learned a lot and probably wouldn't have got so interested in this space if the leak didn't happen.
On the other side of the coin, they've distracted a huge amount of attention from OpenAI and have open source optimisations appearing for every platform they could ever consider running it on, for no extra expense. If it was a deliberate leak, it was a good idea.
Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
#289I've spent an embarassing amount of time since the llamas leaked playing with them, the tools to run them, and writing wrappers for them. They are technically alternatives in the sense that they're incomparably better chat bots than anything in the past. But at least for the 30B and under versions (65B is too big for me to run), no matter what fine tuning is done (alpaca, gpt4all, vicuna, etc), the llamas themselves…
The difference between 3.5 and 4 is gigantic even in my fairly limited experience. I gave them both some common sense tests and this one stuck out to me. Q: A glass door has ‘push’ written on it in mirror writing. Should you push or pull it GPT-3.5: If the word "push" is written in mirror writing on a glass door, you should push the door to open it GPT-4: Since the word "push" is written in mirror writing, it suggest…
> Mirror writing is when words are spelled backwards; this can be done to make the text more visible for people approaching the door from the opposite side. However, since the action required is to open the door, the correct direction would be 'pull' rather than 'push'.
full prompt:
Write a response that appropriately answers the following question, provide your reasoning.
### Instruction:
A glass door has ‘push’ written on it in mirror writing. Should you push or pull it
### Response: