Live data from Hacker News

DBRX: A new open LLM

databricks.com

271–280 of 360 posts

Re: DBRX: A new open LLM

#271
post #249

Earlier quoted context omitted.

What? Are you asking if the framework automatically quantizes/prunes the model on the fly? Or are you suggesting the LLM itself should realize it's too big to run, and prune/quantize itself? Your references to "intelligent" almost leads me to the conclusion that you think the LLM should prune itself. Not only is this a chicken and egg problem, but LLMs are statistical models, they aren't inherently self bootstraping.

I realize that, but I do think it's doable to bootstrap it on a cluster and teach itself to self-prune, and surprised nobody is actively working on this. I hate software that complains (about dependencies, resources) when you try to run it and I think that should be one of the first use cases for LLMs to get L5 autonomous software installation and execution.

Make your dreams a reality!

Re: DBRX: A new open LLM

#272
post #240

I would note the actual leading models right now (IMO) are: - Miqu 70B (General Chat) - Deepseed 33B (Coding) - Yi 34B (for chat over 32K context) And of course, there are finetunes of all these. And there are some others in the 34B-70B range I have not tried (and some I have tried, like Qwen, which I was not impressed with). Point being that Llama 70B, Mixtral and Grok as seen in the charts are not what I would call…

Miqu is a leaked model -- no license is provided to use it. Yi 34B doesn't allow commercial use. Deepseed 33B isn't much good at stuff outside of coding. So it's fair to say that DBRX is the leading general purpose model that can be used commercially.

[dead]

Re: DBRX: A new open LLM

#273
post #213
post #211

Earlier quoted context omitted.

That question was just an example (Lorem ipsum), it was easy to copy paste to demo the local LLM, I didn't intend to provide more context to the discussion. I ordered a 2nd 3090, which has 24GB VRAM. Funny how it was $2.6k 3 years ago and now is $600. You can probuild a decent AI local machine for around $1000.

https://howmuch.one/product/average-nvidia-geforce-rtx-3090-... you are right there is a huge drop in price

New it's hard to find, but the 2nd hand market is filled with them.

Re: DBRX: A new open LLM

#274
post #41

Earlier quoted context omitted.

I have 128GB, but something is weird with Ollama. Even though for the Ollama Docker I only allow 90GB, it ends up using 128GB/128GB, so the system because very slow (mouse freezes).

What docker flags are you running?

None? The default ones from their docs.

The Docker also shows minimal usage for the ollama server which is also strange.

Re: DBRX: A new open LLM

#275
post #240

I would note the actual leading models right now (IMO) are: - Miqu 70B (General Chat) - Deepseed 33B (Coding) - Yi 34B (for chat over 32K context) And of course, there are finetunes of all these. And there are some others in the 34B-70B range I have not tried (and some I have tried, like Qwen, which I was not impressed with). Point being that Llama 70B, Mixtral and Grok as seen in the charts are not what I would call…

Miqu is a leaked model -- no license is provided to use it. Yi 34B doesn't allow commercial use. Deepseed 33B isn't much good at stuff outside of coding. So it's fair to say that DBRX is the leading general purpose model that can be used commercially.

This only applies to projects whose authors seek to comply with the whims of a particular jurisdiction.

Surely there are plenty of project prospects - even commercial in nature - which don't have this limitation.

Re: DBRX: A new open LLM

#276

this proves that all llm models converge to a certain point when trained on the same data. ie, there is really no differentiation between one model or the other. Claims about out-performance on tasks are just that, claims. the next iteration of llama or mixtral will converge. LLMs seem to evolve like linux/windows or ios/android with not much differentiation in the foundation models.

Of course, part of this is that a lot of LLMs are now being trained on data that is itself LLM-generated...

Re: DBRX: A new open LLM

#277
post #240

I would note the actual leading models right now (IMO) are: - Miqu 70B (General Chat) - Deepseed 33B (Coding) - Yi 34B (for chat over 32K context) And of course, there are finetunes of all these. And there are some others in the 34B-70B range I have not tried (and some I have tried, like Qwen, which I was not impressed with). Point being that Llama 70B, Mixtral and Grok as seen in the charts are not what I would call…

Miqu is a leaked model -- no license is provided to use it. Yi 34B doesn't allow commercial use. Deepseed 33B isn't much good at stuff outside of coding. So it's fair to say that DBRX is the leading general purpose model that can be used commercially.

[flagged]

Re: DBRX: A new open LLM

#278
post #240

Earlier quoted context omitted.

Miqu is a leaked model -- no license is provided to use it. Yi 34B doesn't allow commercial use. Deepseed 33B isn't much good at stuff outside of coding. So it's fair to say that DBRX is the leading general purpose model that can be used commercially.

[flagged]

You're being downvoted because everyone in here is looking to profiteer the same way one day.

Re: DBRX: A new open LLM

#279
post #240

I would note the actual leading models right now (IMO) are: - Miqu 70B (General Chat) - Deepseed 33B (Coding) - Yi 34B (for chat over 32K context) And of course, there are finetunes of all these. And there are some others in the 34B-70B range I have not tried (and some I have tried, like Qwen, which I was not impressed with). Point being that Llama 70B, Mixtral and Grok as seen in the charts are not what I would call…

Miqu is a leaked model -- no license is provided to use it. Yi 34B doesn't allow commercial use. Deepseed 33B isn't much good at stuff outside of coding. So it's fair to say that DBRX is the leading general purpose model that can be used commercially.

Model weights are just constants in a mathematical equation, they aren’t copyrightable. It’s questionable whether licenses to use them only for certain purposes are even enforceable. No human wrote the weights so they aren’t a work of art/authorship by a human. Just don’t use their services, use the weights at home on your machines so you don’t bypass some TOS.

Re: DBRX: A new open LLM

#280

The approval on the base model is not feeling very open. Plenty of people still waiting on a chance to download it, where as the instruct model was an instant approval. The base model is more interesting to me for finetuning.

The license allows to reproduce/distribute/copy, so I'm a little surprised there's an approval process at all.

Yeah it's kind of weird, I'll assume for now they're just busy, but I'd be lying if my gut didn't immediately say it's kind of sketchy.
Post reply on HN