Live data from Hacker News

Llama 2

ai.meta.com

671–680 of 860 posts

Re: Llama 2

#671

Earlier quoted context omitted.

> The brain is the result of maximizing biological objective functions. That's not how evolution works at all .

a mutation happens and if that mutation succeeds in ensuring survival, it stays and then spreads. Reproduce is a function evolution maximizes for. Not intentionally sure but that's irrelevant. The whole point of artificial neural networks is that they teach themselves. They get an answer wrong, numbers shift and if those numbers help the next instance they stay or shift as needed. There's no intentionality in the shi…

[deleted]

Re: Llama 2

#673
post #65

Would really want to see some benchmarks against ChatGPT / GPT-4. The improvements in the given benchmarks for the larger models (Llama v1 65B and Llama v2 70B) are not huge, but hard to know if still make a difference for many common use cases.

It would be nice to see 6 of them trained for different purposes by combining 5 of their outputs together and 1 trained to summarize for the most complete and correct output. If we are to trust the leaks about GPT-4, this may be a more fair comparison, even if it is only ~10-20% of the size or so.

Isn't that essentially beam sampling?

Re: Llama 2

#674

Earlier quoted context omitted.

Sparse MoE models are neither new nor secret. The only reason you haven't seen much use of them for LLMs is because they would typically well underperform their dense counterparts. Until this paper ( https://arxiv.org/abs/2305.14705 ) indicated they apparently benefit far more from Instruct tuning than dense models, it was mostly a "good on paper" kind of thing. In the paper, you can see the underperformance i'm talk…

This paper came out well after GPT-4, so apparently this was indeed a secret before then.

Is there a difference here between a secret and an unknown? It may well be that some researcher / comp engineer had an idea, tried it out, realized it was incredibly powerful, implemented it for real this time and then published findings after they were sure of it?

I'm more of a mechanical engineering adjacent professional than a programmer and only follow AI developments loosely

Re: Llama 2

#675

Earlier quoted context omitted.

>It either swims or it doesnt Correct, it swims. >A drowning hippo isn't going to wish itself to float. A drowning hippo probably wishes it can float, much like a drowning person wishes they can float.

Well, people can float. Also people can swim, so even if they were super muscular and lean and this made them incapable of floating (I don’t know if that happens), they could swim if they knew how. It sounds like hippos in deep water are incapable of swimming to the top. Based on what I am reading in this thread, they would simply sink. Humans, properly instructed, can avoid this by swimming.

A properly instructed hippo would stay out of the deep end

Re: Llama 2

#676
I keep getting this - been trying sporadically over the past couple hours. Anyone else hit this and any way to work around this

Resolving download.llamameta.net (download.llamameta.net)... 108.138.94.71, 108.138.94.95, 108.138.94.120, ... Connecting to download.llamameta.net (download.llamameta.net)|108.138.94.71|:443... connected. HTTP request sent, awaiting response... 403 Forbidden 2023-07-18 18:02:19 ERROR 403: Forbidden.

Re: Llama 2

#677

Earlier quoted context omitted.

The 70B Llama2 model ties in with 173B ChatGPT-0301 model. The GPT-4 still stands unchallenged.

Source on the 173B parameters?

The wikipedia article for GPT-4 has this as its source: https://the-decoder.com/gpt-4-architecture-datasets-costs-an...

Re: Llama 2

#678
post #558
post #515

Here are some benchmarks, excellent to see that an open model is approaching (and in some areas surpassing) GPT-3.5! AI2 Reasoning Challenge (25-shot) - a set of grade-school science questions. - Llama 1 (llama-65b): 57.6 - LLama 2 (llama-2-70b-chat-hf): 64.6 - GPT-3.5: 85.2 - GPT-4: 96.3 HellaSwag (10-shot) - a test of commonsense inference, which is easy for humans (~95%) but challenging for SOTA models. - Llama 1:…

Is it possible that some LLM’s are trained on these benchmarks? Which would mean they’re overfitting and are incorrectly ranked? Or am I misunderstanding these benchmarks?…

Unfortunately, Goodhart's law applies on most kind of tests

> Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.

Re: Llama 2

#679

Earlier quoted context omitted.

If you recreated those works from memory yeah you would be subject to copyright. There's a whole set of rules around fair use and derivative work.

Where is AI guilty of reproducing Star Wars verbatim, then? If the AI has seen Star Wars and that's enough to find it liable, then you should be too. If the AI has seen Star Wars to understand science fiction and modern culture, then it's no different from us or any other artist.

If a human recites large chunks of Star Wars verbatim, and then sells that copy as a service, thats certainly enough to find the person liable.

YouTube zaps videos that contain too much copyrighted stuff for this very reason.

Re: Llama 2

#680
post #558
post #515

Here are some benchmarks, excellent to see that an open model is approaching (and in some areas surpassing) GPT-3.5! AI2 Reasoning Challenge (25-shot) - a set of grade-school science questions. - Llama 1 (llama-65b): 57.6 - LLama 2 (llama-2-70b-chat-hf): 64.6 - GPT-3.5: 85.2 - GPT-4: 96.3 HellaSwag (10-shot) - a test of commonsense inference, which is easy for humans (~95%) but challenging for SOTA models. - Llama 1:…

Is it possible that some LLM’s are trained on these benchmarks? Which would mean they’re overfitting and are incorrectly ranked? Or am I misunderstanding these benchmarks?…

Presented with no comment :) https://twitter.com/chhillee/status/1635790330854526981?s=46...
Post reply on HN