Live data from Hacker News

Llama 2

ai.meta.com

571–580 of 860 posts

Re: Llama 2

#571
post #558
post #515

Here are some benchmarks, excellent to see that an open model is approaching (and in some areas surpassing) GPT-3.5! AI2 Reasoning Challenge (25-shot) - a set of grade-school science questions. - Llama 1 (llama-65b): 57.6 - LLama 2 (llama-2-70b-chat-hf): 64.6 - GPT-3.5: 85.2 - GPT-4: 96.3 HellaSwag (10-shot) - a test of commonsense inference, which is easy for humans (~95%) but challenging for SOTA models. - Llama 1:…

Is it possible that some LLM’s are trained on these benchmarks? Which would mean they’re overfitting and are incorrectly ranked? Or am I misunderstanding these benchmarks?…

It would be a bit of a scandal, and IMO too much hassle to sneak in. These models are trained on massive amounts of text - specifically anticipating which metrics people will care about and generating synthetic data just for them seems extra.

But not an expert or OP!

Re: Llama 2

#572

Earlier quoted context omitted.

Good to see these results, thanks for posting. I wonder if GPT-4's dominance is due to some secret sauce or if its just the first mover advantage and Llama will be there soon.

It's just scale. But scale that comes with more than an order of magnitude more expense than the Llama models. I don't see anyone training such a model and releasing it for free anytime soon

I thought it was revealed to be fundamentally ensemblamatic in a way the others weren’t? Using “experts” I think? Seems like it would meet the bar for “secret sauce” to me

Re: Llama 2

#573
post #60

Earlier quoted context omitted.

Google has far better models than llama based models. They just simply don't put them facing the public. It is pretty ridiculous that they essentially just set a marketing team with no programming experience to write Bard, but that shouldn't fool anyone into believing they don't have capable models in Google. If Deepmind were to actually provide what they have in some usable form, it would likely be quite good. Despi…

Source?

https://github.com/facebookresearch/llama/blob/main/LICENSE#...

Re: Llama 2

#574
post #558
post #515

Here are some benchmarks, excellent to see that an open model is approaching (and in some areas surpassing) GPT-3.5! AI2 Reasoning Challenge (25-shot) - a set of grade-school science questions. - Llama 1 (llama-65b): 57.6 - LLama 2 (llama-2-70b-chat-hf): 64.6 - GPT-3.5: 85.2 - GPT-4: 96.3 HellaSwag (10-shot) - a test of commonsense inference, which is easy for humans (~95%) but challenging for SOTA models. - Llama 1:…

Is it possible that some LLM’s are trained on these benchmarks? Which would mean they’re overfitting and are incorrectly ranked? Or am I misunderstanding these benchmarks?…

Test leakage is not impossible for some benchmarks. But researchers try to avoid/mitigate that as much as possible for obvious reasons.

Re: Llama 2

#575

Earlier quoted context omitted.

That's just stupid talk. It either swims or it doesnt. A drowning hippo isn't going to wish itself to float.

>It either swims or it doesnt Correct, it swims. >A drowning hippo isn't going to wish itself to float. A drowning hippo probably wishes it can float, much like a drowning person wishes they can float.

Well, people can float. Also people can swim, so even if they were super muscular and lean and this made them incapable of floating (I don’t know if that happens), they could swim if they knew how. It sounds like hippos in deep water are incapable of swimming to the top. Based on what I am reading in this thread, they would simply sink. Humans, properly instructed, can avoid this by swimming.

Re: Llama 2

#577
Offtopic, I know. But I was wondering why the site loaded slowly on my phone. They're using images for everything: benchmark tables (rendered from HTML?), background gradients. One gradient is a 2MB PNG.

Re: Llama 2

#578

Earlier quoted context omitted.

If so then that means the training objective is wrong because admitting you do not know something is much more a hallmark of intelligence than any attempt to 'hallucinate' (I don't like that word, I prefer 'make up') an answer.

I guess the brains objective is wrong then seeing how much it's willing to fabricate sense data, memories and rationales when convenient

The brain wasn't designed.

Re: Llama 2

#579
post #60

Earlier quoted context omitted.

Google has far better models than llama based models. They just simply don't put them facing the public. It is pretty ridiculous that they essentially just set a marketing team with no programming experience to write Bard, but that shouldn't fool anyone into believing they don't have capable models in Google. If Deepmind were to actually provide what they have in some usable form, it would likely be quite good. Despi…

Google's LLMs are all vaporware. No one's ever seen them. They're supposedly mind-blowing but when they are released they always sound like lobotomized monkeys. All the AlphaGo/AlphaFold stuff is very cool, but since no one has seen their LLMs this is about as convincing as my claiming I've donated billions to charity.

I can assure you Google BERT isn't vaporware.

It was probably a challenge to integrate it into search, but they did that.

So your assertion has been refuted based on your use of "all", at the very least.

Re: Llama 2

#580
post #515

Here are some benchmarks, excellent to see that an open model is approaching (and in some areas surpassing) GPT-3.5! AI2 Reasoning Challenge (25-shot) - a set of grade-school science questions. - Llama 1 (llama-65b): 57.6 - LLama 2 (llama-2-70b-chat-hf): 64.6 - GPT-3.5: 85.2 - GPT-4: 96.3 HellaSwag (10-shot) - a test of commonsense inference, which is easy for humans (~95%) but challenging for SOTA models. - Llama 1:…

Good to see these results, thanks for posting. I wonder if GPT-4's dominance is due to some secret sauce or if its just the first mover advantage and Llama will be there soon.

GPT4 is rumored to have 1.7T parameters, Llama 2 70B.
Post reply on HN