Live data from Hacker News

Meta Llama 3

llama.meta.com

921–930 of 965 posts

Re: Meta Llama 3

#921
post #4

They've got a console for it as well, https://www.meta.ai/ And announcing a lot of integration across the Meta product suite, https://about.fb.com/news/2024/04/meta-ai-assistant-built-wi... Neglected to include comparisons against GPT-4-Turbo or Claude Opus, so I guess it's far from being a frontier model. We'll see how it fares in the LLM Arena.

Why does Meta embed a 3.5MB animated GIF (https://about.fb.com/wp-content/uploads/2024/04/Meta-AI-Expa...) on their announcement post instead of much smaller animated WebP/APNG/MP4 file? They should care about users with low bandwidth and limited data plan.

Re: Meta Llama 3

#922

Earlier quoted context omitted.

Lex walked so that Dwarkesh could run. He runs the best AI podcast around right now, by a long shot.

I don't know Dwarkesh but I despise Lex Fridman. I don't know how a man that lacks the barest modicum of charisma has propelled himself to helming a high-profile, successful podcast. It's not like he tends to express interesting or original thoughts to make up for his paucity of presence. It's bizarre. Maybe I'll check out Dwarkesh, but even seeing him mentioned him in the same breath as Fridman gives me pause ...

I listen to Lex relatively often. I think he often has enough specialized knowledge to keep up at least somewhat with guests. His most recent interview of the Egyptian comedian (not a funny interview) on Palestine was really profound, as in one of the best podcasts I’ve ever listened to.

Early on I got really fed up with him when I discovered him. Like his first interview with mark zuckerberg where he asks him multiple times to basically say his life is worthless, his huge simping to Elon musks, asking empty questions repeatedly, and being jealous of Mr Beast.

But yeah for whatever reason lately I’ve dug his podcast a lot. Those less good interviews were from a couple years ago. Though I wish he didn’t obsess so much about twitter

Re: Meta Llama 3

#923

Earlier quoted context omitted.

I agree that it is the best AI podcast. I do have a few gripes though, which might just be from personal preference. A lot of the time the language used by both the host and the guests is unnecessarily obtuse. Also the host is biased towards being optimistic about LLMs leading to AGI, and so he doesn't probe guests deep enough about that, more than just asking something along the lines of "Do you think next token pre…

There's a difference to being a good chatshow/podcast host and a journalist holding someone's feet to the fire! Dwarkesh is excellent at what he does - lots of research beforehand (which is how he lands these great guests), but then lets the guest do most of the talking, and encourages them to expand on what they are saying. It you are critisizing the guest or giving them too much push back, then they are going to cl…

I decided to listen to a Dwarkesh episode as a result of this thread. I chose the Eliezer Yudkowsky episode. After 90 minutes, Dwarkesh is raising one of the same 3 objections for the n-teenth time, instead of leading the conversation in an interesting direction. If his other AI episodes are in the vein as other comments describe, then this does seem to be plain old positive AGI optimism bias rather than some special interview technique. In addition, he's very ill-prepared in that he doesn't seem to have attempted to understand the reasons some people have for believing AGI to be a threat.

On the other hand, Yudkowsky was a terrible guest, in terms of his public speaking skills. He came across as combative. His answers were terse and he spent little time on background information or otherwise making an effort to explain his reasoning in a way more digestible for a general audience.

Re: Meta Llama 3

#924
post #4

They've got a console for it as well, https://www.meta.ai/ And announcing a lot of integration across the Meta product suite, https://about.fb.com/news/2024/04/meta-ai-assistant-built-wi... Neglected to include comparisons against GPT-4-Turbo or Claude Opus, so I guess it's far from being a frontier model. We'll see how it fares in the LLM Arena.

Losers & Winners from Llama-3-400B Matching 'Claude 3 Opus' etc.. Losers: - Nvidia Stock : lid on GPU growth in the coming year or two as "Nation states" use Llama-3/Llama-4 instead spending $$$ on GPU for own models, same goes with big corporations. - OpenAI & Sam: hard to raise speculated $100 Billion, Given GPT-4/GPT-5 advances are visible now. - Google : diminished AI superiority posture Winners: - AMD, intel: th…

If anything a capable open source model is good for Nvidia, not commenting on their share price but business of course.

Better open models lower the barrier to build products and drive the price down, more options at cheaper prices which means bigger demand for GPUs and Cloud. More of what the end customers pay for goes to inference and not IP/training of proprietary models

Re: Meta Llama 3

#925

Earlier quoted context omitted.

Also, being open source adds phenomenal value for Meta: 1. It attracts the world's best academic talent, who deeply want their work shared. AI experts can join any company, so ones which commit to open AI have a huge advantage. 2. Having armies of SWEs contributing millions of free labor hours to test/fix/improve/expand your stuff is incredible. 3. The industry standardizes around their tech, driving down costs and d…

Commoditize Your Complements: https://gwern.net/complement

No need to quote the arrogant clown on that one, Spolski coined the concept:

https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/

Re: Meta Llama 3

#926
post #242

Earlier quoted context omitted.

https://ai.meta.com/blog/meta-llama-3/ has it about a third of the way down. It's a little bit better on every benchmark than Mixtral 8x22B (according to Meta).

Oh cool! But at the cost of twice the VRAM and only having 1/8th of the context, I suppose?

Llama 3 70B takes half the VRAM as Mixtral 8x22B. But it does need almost twice the FLOPS/bandwidth. Yes, Llama's context is smaller although that should be fixable in the near future. Another thing is that Llama is English-focused while Mixtral is more multilingual.

Re: Meta Llama 3

#927

Earlier quoted context omitted.

"Censored" is the word that you're looking for, and is generally what you see when these models are discussed on Reddit etc. Not to worry - uncensored finetunes will be coming shortly.

You can't really take out the censorship. You can strengthen pathways which work around the damage, but the damage is still there.

If the model doesn't refuse to produce output, it's not censored anymore for any practical purpose. It doesn't really matter if there are "censorship neurons" inside that are routed around.

Sure, it would be nice if we didn't have to do that so that the model could actually spent its full capacity on something useful. But that's a different issue even if the root cause is the same.

Re: Meta Llama 3

#928

Earlier quoted context omitted.

I also disagree on Google... Google's business is largely not predicated on AI the way everyone else is. Sure they hope it's a driver of growth, but if the entire LLM industry disappeared, they'd be fine. Google doesn't need AI "Superiority", they need "good enough" to prevent the masses from product switching. If the entire world is saturated in AI, then it no longer becomes a differentiator to drive switching. And…

Google’s play is not really in AI imo, it’s in the the fact that their custom silicon allows them to run models cheaply. Models are pretty much fungible at this point if you’re not trying to do any LoRAs or fine tunes.

There's still no other model on par with GPT-4. Not even close.

Re: Meta Llama 3

#929
post #808

Earlier quoted context omitted.

>AMD, intel: these companies can focus on Chips for AI Inference No real evidence either can pull that off in any meaningful timeline, look how badly they neglected this type of computing the past 15 years.

AMD is already competitive on inference

Their problem is that the ecosystem is still very CUDA-centric as a whole.

Re: Meta Llama 3

#930
post #527

Earlier quoted context omitted.

EU actually has the opposite of draconian privacy laws. It's more that meta doesn't have a business model if they don't intrude on your privacy

They just said laws, not privacy - the EU has introduced the "world's first comprehensive AI law". Even if it doesn't stop release of these models, it might be enough that the lawyers need extra time to review and sign off that it can be used without Meta getting one of those "7% of worldwide revenue" type fines the EU is fond of. [0] https://www.europarl.europa.eu/topics/en/article/20230601STO...

Am I reading that right? It sounds like they’re outlawing advertising (“Cognitive behavioural manipulation of people”), credit scores (“classifying people based on behaviour, socio-economic status or personal characteristics”) and fingerprint/facial recognition for phone unlocking etc. (“Biometric identification and categorisation of people”)

Maybe they mean specific uses of these things in a centralised manner but the way it’s written makes it sound incredibly broad.

Post reply on HN