Live data from Hacker News

Meta Llama 3

llama.meta.com

861–870 of 965 posts

Re: Meta Llama 3

#861
post #552

Earlier quoted context omitted.

Those numbers are for the original GPT-4 (Mar 2023). Current GPT-4-Turbo (Apr 2024) is better: Llama 3 GPT-4 GPT-4-Turbo* (Apr 2024) MMLU 86.1 86.4 86.7 DROP 83.5 80.9 86.0 MATH 57.8 52.9 73.4 HumEv 84.1 74.4 88.2 *using API prompt: https://github.com/openai/simple-evals

I find it somewhat interesting that there is a common perception about GPT-4 at release being actually smart, but that it got gradually nerfed for speed with turbo, which is better tuned but doesn't exhibit intelligence like the original. There were times when I felt that too, but nowadays I predominantly use turbo. It's probably because turbo is faster and cheaper, but in lmsys turbo has 100 elo higher than original…

i think it might just be the subjective feelings (GPT-4-turbo being dumber) - the joy is always stronger when you first taste it, and the joy decays as you get used to it and the bar raises ever since.

Re: Meta Llama 3

#862

Earlier quoted context omitted.

His engineering mindset made him blind to the fact the metaverse was a product that nobody wanted or needed. In one of the Fridman interviews, he goes on and on about all the cool technical challenges involved in making the metaverse work. But when Fridman asked him what he likes to do in his spare time, it was all things that you could precisely not do in the metaverse. It was baffling to me that he failed to connec…

I don't think that was the issue. VRChat was basically the same idea but done in a more appealing way and it was (still is) wildly popular.

VRChat is more popular, but it doesn’t mean that copying their approaches would be the move.

For all we know, VRChat as a concept of that kind is a local maximum, and imo it wont scale well to genpop. Not claiming this as an objective fact, but as a hypothesis that I personally believe to be very likely truthful. Think of it as a dead branch of evolution, where if you want to go further than that local maximum, you gotta break out of it using an entirely different approach.

I like VRChat, but thinking that a random person living in the mainstream who isnt into that type of geeky online stuff is gonna be convinced of VRChat being the ultimate metaverse experience is just foolish.

At that point, your choices are: (1) build a VRChat clone and hit that same local maximum but slightly higher at best or (2) develop something entirely different to get out of that local maximum, but risk failing (since it is a totally novel thing) and coming short of being at least as successful as VRChat. Zuck took the second option, and I respect that.

Just making a VRChat Meta Edition clone would imo give Meta much better numbers in the short-term (than their failed Meta Horizons did), but imo long-term that approach would lead them nowhere. And it seems like Meta is more interested in capturing the first-mover (into the mainstream) advantage heavy.

And honestly, I think it is better off this way. Just like if someone is making yet another group chat, i would prefer they went balls to the wall, tried to rethink things from scratch, and made a group chat app that is unlike any other ones out there. Could all of their novel approaches fail? Yes, much more likely than if they made another slack clone with a different color scheme. But the important part is, it also has a much higher chance to get the state of their niche oit of the local maximum.

Examples: Twitter could’ve been just another blog aggregator, Tesla could’ve been just another gas-powered Lotus Elise (with the original roadsters literally being just their custom internals slotted into a Lotus body), Microsoft would’ve been stuck with MS-DOS and not went into the “app as the main OS” thing (which is what they did with Windows).

Apple would’ve been relegated to a legacy of Apple II and iPod (with a dash of macbook relevancy), and rememebered as the company that made this ultra popular mp3 player before that whole niche died. Airpods (that everyone laughed at initially and lauded as an impractical pretentious purchase) are massive now, with every holdout that I personally know who finally got them recently going “i cannot believe how convenient it is, i should’ve gotten them earlier”, but it was also a similar “who needs this, they are solving a problem nobody has, everyone prefers wired with tons of better options” take[0].

If you want to get out of a perceived local maximum and break into the mainstream, you gotta try brand new approaches that would likely fail. Going “omg cannot even beat that existing competitor that’s been running for years” is kinda pointless in this case, because competing with them directly by making just a better and more successful clone of their product was never the goal. I don’t doubt even for a second that if Meta tried that, they would’ve likely accomplished it.

And for the naysayers who don’t see Meta ever breaking things out of a local maximum, just look at the Oculus Quest line. Everyone was laughing at them initially for going with the standalone device approach, but Quest has become a massive hit, with tons of people of all kinds buying it (not just people with massive gaming rigs).

0. And yes, removal of the audiojack somewhat speeded up the adoption, but I just used an adapter with zero discomfort for a year or two until i got airpods myself (and would’ve still continued using the adapter if I just didnt flatout preferred airpods in general).

Re: Meta Llama 3

#863

Earlier quoted context omitted.

They didn't compare against the best models because they were trying to do "in class" comparisons, and the 70B model is in the same class as Sonnet (which they do compare against) and GPT3.5 (which is much worse than sonnet). If they're beating sonnet that means they're going to be within stabbing distance of opus and gpt4 for most tasks, with the only major difference probably arising in extremely difficult reasonin…

Llama is open weight, not open source. They don’t release all the things you need to reproduce their weights.

Which large model projects are open source in that sense? That its full source code including training material is published.

Re: Meta Llama 3

#864
post #112

Earlier quoted context omitted.

The bottom of https://ai.meta.com/blog/meta-llama-3/ has in-progress results for the 400B model as well. Looks like it's not quite there yet. Llama 3 400B Base / Instruct MMLU 84.8 86.1 GPQA - 48.0 MATH - 57.8 HumanEval - 84.1 DROP 83.5 -

Not quite there yet, but very close and not done training! It's quite plausible that this model could be state of the art over GPT-4 in some domains when it finishes training, unless GPT-5 comes out first. Although 400B will be pretty much out of reach for any PC to run locally, it will still be exciting to have a GPT-4 level model in the open for research so people can try quantizing, pruning, distilling, and other…

There are rumors about an upcoming M3 or M4 Extreme chip... which would certainly have enough RAM, and probably a 1600-2000 GB/s bandwidth.

Still wouldn't be super performant AFA token gen, ~4-6 per second, but certainly runnable.

Of course by the time that lands in 6-12 months we'll probably have a 70-100G model that is similarly performant.

Re: Meta Llama 3

#865

I imagine it's a given at this point, but I figured it was worth noting that it seems they trained this using OpenAI outputs. Using meta.ai to test the model, it gave me a link to a google search when questioned about a relatively current event. When I expressed surprise that it could access the internet it told me it did so via Bing. I asked it to clarify why it said Bing, when it gave me an actual link to a google…

You really should know better than to interrogate an LLM about itself. They do not have self-awareness and will readily hallucinate. "Meta also announced a partnership with Google to include its real-time search results in the assistant's responses, supplementing an existing arrangement with Microsoft's Bing search engine." from https://www.reuters.com/technology/meta-releases-early-versi...

Appreciate the additional information!

Re: Meta Llama 3

#866

Earlier quoted context omitted.

You can build a machine that can run 70b models at great TpS speeds for around 30-60k. That same machine could almost certainly run a 400b model with "useable" speeds. Obviously much slower than current ChatGPT speeds but still, that kind of machine is well within the means of wealthy hobbyists/highly compensated SWEs and small firms.

I just tested llama3:70b with ollama on my old AMD ThreadRipper Pro 3965WX workstation (16-core Zen4 with 8 DDR4 mem channels), with a single RTX 4090. Got 3.5-4 tokens/s, GPU compute was <20% busy (~90W) and the 16 CPU cores / 32 threads were about 50% busy.

And that’s not quantized at all, correct?

If so, then the parent comment’s sentiment holds true…. Exciting stuff.

Re: Meta Llama 3

#867
Llama 3 70B has debuted on the famous LMSYS chatbot arena leaderboard at position number 5, tied with Claude 2 Sonnet, Bard (Gemini Pro), and Command R+, ahead of Claude 2 Haiku and older versions of GPT-4.

The score still has a large uncertainty so it will take a while to determine the exact ranking and things may change.

Llama 3 8B is at #12 tied with Claude 1, Mixtral 8x22B, and Qwen-1.5-72B.

These rankings seem very impressive to me, on the most trusted benchmark around! Check the latest updates at https://arena.lmsys.org/

Edit: On the English-only leaderboard Llama 3 70B is doing even better, hovering at the very top with GPT-4 and Claude Opus. Very impressive! People seem to be saying that Llama 3's safety tuning is much less severe than before so my speculation is that this is due to reduced refusal of prompts more than increased knowledge or reasoning, given the eval scores. But still, a real and useful improvement! At this rate, the 400B is practically guaranteed to dominate.

Re: Meta Llama 3

#868
post #820

Earlier quoted context omitted.

I also disagree on Google... Google's business is largely not predicated on AI the way everyone else is. Sure they hope it's a driver of growth, but if the entire LLM industry disappeared, they'd be fine. Google doesn't need AI "Superiority", they need "good enough" to prevent the masses from product switching. If the entire world is saturated in AI, then it no longer becomes a differentiator to drive switching. And…

AI is taking marketshare from search slowly. More and more people will go to the AI to find things and not a search bar. It will be a crisis for Google in 5-10 years.

Source?

Re: Meta Llama 3

#869

Earlier quoted context omitted.

I'm not OP, but George Hotz said in his lex friedman podcast a while back that it was an MoE of 8 250B. subtract out duplication of attention nodes, and you get something right around 1.8T

I'm pretty sure he suggested it was a 16 way 110 MoE

The exact quote: "Sam Altman won’t tell you that GPT 4 has 220 billion parameters and is a 16 way mixture model with eight sets of weights."
Post reply on HN