Live data from Hacker News

DBRX: A new open LLM

databricks.com

211–220 of 360 posts

Re: DBRX: A new open LLM

#211
post #192
post #39

Earlier quoted context omitted.

> No, it can't run at all. https://s3.amazonaws.com/i.snag.gy/ae82Ym.jpg EDIT: This was ran on a 1080ti + 5900x. Initial generation takes around 10-30seconds (like it has to upload the model to GPU), but then it starts answering immediately, at around 3 words per second.

this is some new flex to debate online: copying and pasting the other sides argument and waiting for your local LLM to explain why they are wrong. how much is your hardware at today's value? what are the specs? that is impressive even though its 3 words per second. if you want to bump it up to 30, do you then 10x your current hardware cost?

That question was just an example (Lorem ipsum), it was easy to copy paste to demo the local LLM, I didn't intend to provide more context to the discussion.

I ordered a 2nd 3090, which has 24GB VRAM. Funny how it was $2.6k 3 years ago and now is $600.

You can probuild a decent AI local machine for around $1000.

Re: DBRX: A new open LLM

#212

Earlier quoted context omitted.

It's even possible they converge when trained on different data, if they are learning some underlying representation. There was recent research on face generation where they trained two models by splitting one training set in two without overlap, and got the two models to generate similar faces for similar conditioning, even though each model hadn't seen anything that the other model had.

That sounds unsurprising? Like if you take any set of numbers, randomly split it in two, then calculate the average of each half... it's not surprising that they'll be almost the same. If you took two different training sets then it would be more surprising. Or am I misunderstanding what you mean?

It doesn't really matter whether you do this experiment with two training sets created independently or one training set split in half. As long as both are representative of the underlying population, you would get roughly the same results. In the case of human faces, as long as the faces are drawn from roughly similar population distributions (age, race, sex), you'll get similar results. There's only so much variation in human faces.

If the populations are different, then you'll just get two models that have representations of the two different populations. For example, if you trained a model on a sample of all old people and separately on a sample of all young people, obviously those would not be expected to converge, because they're not drawing from the same population.

But that experiment of splitting one training set in half does tell you something: the model is building some sort of representation of the underlying distribution, not just overfitting and spitting out chunks of copy-pasted faces stitched together.

Re: DBRX: A new open LLM

#213
post #211
post #192

Earlier quoted context omitted.

this is some new flex to debate online: copying and pasting the other sides argument and waiting for your local LLM to explain why they are wrong. how much is your hardware at today's value? what are the specs? that is impressive even though its 3 words per second. if you want to bump it up to 30, do you then 10x your current hardware cost?

That question was just an example (Lorem ipsum), it was easy to copy paste to demo the local LLM, I didn't intend to provide more context to the discussion. I ordered a 2nd 3090, which has 24GB VRAM. Funny how it was $2.6k 3 years ago and now is $600. You can probuild a decent AI local machine for around $1000.

https://howmuch.one/product/average-nvidia-geforce-rtx-3090-... you are right there is a huge drop in price

Re: DBRX: A new open LLM

#214
post #107

Earlier quoted context omitted.

Their goal is to always drive enterprise business towards consumption. With AI they need to desperately steer the narrative away from API based services (OpenAI). By training LLMs, they build sales artifacts (stories, references, even accelerators with LLMs themselves) to paint the pictures needed to convince their enterprise customer market that Databricks is the platform for enterprise AI. Their blog details how th…

Thanks! Why do they not focus on hosting other open models then? I suspect other models will soon catch up with their advantages in faster inference and better benchmark results. That said, maybe the advantage is aligned interests: they want customers to use their platforms, so they can keep their models open. In contrast, Mistral removed their commitment to open source as they found a potential path to profitability…

> Why do they not focus on hosting other open models then?

They do host other open models as well (pay-per-token).

Re: DBRX: A new open LLM

#215
post #90

Earlier quoted context omitted.

I don't think it's possible to have an "open training data" model because it would get DMCA'd immediately and open you up to lawsuits from everyone who found their works in the training set. I hope we can fix the legal landscape to enable publicly sharing training data but I can't really judge the companies keeping it a secret today.

> I don't think it's possible to have an "open training data" model because it would get DMCA'd immediately… This isn't a problem because OpenAI says, "training AI models using publicly available internet materials is fair use". /s https://openai.com/blog/openai-and-journalism

I don't think it's that crazy, even if you're sure it's fair use I wouldn't paint a huge target on my back before there's a definite ruling and I doubly wouldn't test the waters of the legality of re-hosting copyrighted content to be downloaded by randos who won't be training models with it.

If they're going to get away with this collecting data and having a legal chain-of-custody so you can actually say it was only used to train models and no one else has access to it goes a long way.

Re: DBRX: A new open LLM

#216

Earlier quoted context omitted.

Looks like someone has got DBRX running on an M2 Ultra already: https://x.com/awnihannun/status/1773024954667184196?s=20

And it appears to be at ~80 GB of RAM via quantisation.

So that would be runnable on a MBP with a M2 Max, but the context window must be quite small, I don’t really find anything under about 4096 that useful

Re: DBRX: A new open LLM

#217
post #188

Earlier quoted context omitted.

Databricks is trying to go all-in on convincing organizations they need to use in-house models, and therefore pay they to provide LLMOps. They're so far into this that their CTO co-authored a borderline dishonest study which got a ton of traction last summer trying to discredit GPT-4: https://arxiv.org/pdf/2307.09009.pdf

What does borderline dishonest mean? I only read the abstract and it seems like such an obvious point I dont see how its contentious

The regression came from poorly parsing the results. I came the conclusion independently, but here's another more detailed takedown: https://www.reddit.com/r/ChatGPT/comments/153xee8/has_chatgp...

Given the conflict of interest and background of Zaharia, it's hard to imagine such an immediately obvious source of error wasn't caught.

Re: DBRX: A new open LLM

#218
post #166

this proves that all llm models converge to a certain point when trained on the same data. ie, there is really no differentiation between one model or the other. Claims about out-performance on tasks are just that, claims. the next iteration of llama or mixtral will converge. LLMs seem to evolve like linux/windows or ios/android with not much differentiation in the foundation models.

There's at least an argument to be made that this is because all the models are heavily trained on GPT-4 outputs (or whatever the SOTA happens to be during training). All those models are, in a way, a product of inbreeding.

But is it the kind of inbreeding that gets you Downs, or the kwisatz haderach?

Re: DBRX: A new open LLM

#219
post #160

Worse than the chart crime of truncating the y axis is putting LLaMa2's Human Eval scores on there and not comparing it to Code Llama Instruct 70b. DBRX still beats Code Llama Instruct's 67.8 but not by that much.

> "On HumanEval, DBRX Instruct even surpasses CodeLLaMA-70B Instruct, a model built explicitly for programming, despite the fact that DBRX Instruct is designed for general-purpose use (70.1% vs. 67.8% on HumanEval as reported by Meta in the CodeLLaMA blog)." To be fair, they do compare to it in the main body of the blog. It's just probably misleading to compare to CodeLLaMA on non coding benchmarks.

Which non-coding benchmark?

Re: DBRX: A new open LLM

#220
The approval on the base model is not feeling very open. Plenty of people still waiting on a chance to download it, where as the instruct model was an instant approval. The base model is more interesting to me for finetuning.
Post reply on HN