Live data from Hacker News

Meta Llama 3

llama.meta.com

951–960 of 965 posts

Re: Meta Llama 3

#951
post #813
post #807

Earlier quoted context omitted.

Yet, people still line up to become veterinaries (and technicians). Which proves my point. > Calling these jobs “cute” or saying the veterinary situation is “fortunate” borders on cruel, [...] Perhaps not the best choice of words, I admit.

> Yet, people still line up to become veterinaries (and technicians). Which proves my point. The informed reality is that the rate of drop out is also huge. Not only from people who leave the course while studying, but also professionals who abandon the field entirely after just a few years of work. Many of them are already suffering in college yet continue due to a sense of necessity or sunk cost and burn themselves…

> So no, it does not prove your point. The one thing it proves is that the public in general is insufficiently informed about what being a veterinary is like.

That doesn't really matter. What would matter is how well informed the people who decide to become a veterinary are.

> They should be paid more and have better conditions [...]

Well, everyone should be treated better and paid better.

> [...] because there’s always another chump down the line.

If they could somehow make the improvements you suggest (but don't specify how), they would lead to even more chumps joining the queue.

(And no, that's not a generalised argument against making people's lives better. If you improve the appeal of non-vet jobs, fewer people will join the vet line.

If you improve the treatment of workers in general, the length of the wanna-be-vet queue, and any other 'job queue' will probably stay roughly the same. But people will be better off.)

Re: Meta Llama 3

#952

Earlier quoted context omitted.

Many disagree. “Not even close” is a strong position to take on this.

It takes less than an hour of conversation with either, giving them a few tasks requiring logical reasoning, to arrive at that conclusion. If that is a strong position, it's only because so many people seem to be buying the common scoreboards wholesale.

That’s very subjective and case dependent. I use local models most often myself with great utility and advocate for giving my companies the choice of using either local models or commercial services/APIs (ChatGPT, GPT-4 API, some Llama derivative, etc.) based on preference. I do not personally find there to be a large gap between the capabilities of commercial models and the fine-tuned 70b or Mixtral models. On the whole, individuals in my companies are mixed in their opinions enough for there to not be any clear consensus on which model/API is best objectively — seems highly preference and task based. This is anecdotal (though the population size is not small), but I think qualitative anec-data is the best we have to judge comparatively for now.

I agree scoreboards are not a highly accurate ranking of model capabilities for a variety of reasons.

Re: Meta Llama 3

#954

Earlier quoted context omitted.

They didn't compare against the best models because they were trying to do "in class" comparisons, and the 70B model is in the same class as Sonnet (which they do compare against) and GPT3.5 (which is much worse than sonnet). If they're beating sonnet that means they're going to be within stabbing distance of opus and gpt4 for most tasks, with the only major difference probably arising in extremely difficult reasonin…

Llama is open weight, not open source. They don’t release all the things you need to reproduce their weights.

Is it really useful to make an LLM open source when it takes millions of $ to train it?

At that scale, open weights with permissive license is much more useful than open source.

Re: Meta Llama 3

#955

Earlier quoted context omitted.

It takes less than an hour of conversation with either, giving them a few tasks requiring logical reasoning, to arrive at that conclusion. If that is a strong position, it's only because so many people seem to be buying the common scoreboards wholesale.

That’s very subjective and case dependent. I use local models most often myself with great utility and advocate for giving my companies the choice of using either local models or commercial services/APIs (ChatGPT, GPT-4 API, some Llama derivative, etc.) based on preference. I do not personally find there to be a large gap between the capabilities of commercial models and the fine-tuned 70b or Mixtral models. On the w…

If you're using them mostly for stuff like data extraction (which seems to be the vast majority of productive use so far), there are many models that are "good enough" and where GPT-4 will not demonstrate meaningful improvements.

It's complicated tasks requiring step by step logical reasoning where GPT-4 is clearly still very much in a league of its own.

Re: Meta Llama 3

#956

Earlier quoted context omitted.

Anecdotally speaking I use google search much less frequently and instead opt for GPT4. This is also what a number of my colleagues are doing as well.

I often use ChatGPT4 for technical info. It's easier then scrolling through pages whet it works. But.. the accuracy is inconsistent, to put it mildly. Sometimes it gets stuck on wrong idea. Interesting how far LLMs can get? Looks like we are close to scale-up limit. It's technically difficult to get bigger models. The way to go probably is to add assisting sub-modules. Examples would be web search, have it already. D…

I expect a 5x improvement before EOY, I think GPT5 will come out.

Re: Meta Llama 3

#957

Earlier quoted context omitted.

Yes how dare different people have different opinions about different people? It's almost as if we all should be a monolithic voice that agrees with you.

The thread was suspiciously positive, like almost exclusive. Your comment adds nothing to the discussion, you're just snarky and nothing else. So get off my back

>Your comment adds nothing to the discussion,

and yours did? This comment, Christian?

>>I swear, this feels like people get paid to write positive stuff about him?

----

>you're just snarky and nothing else

Please re-read your own comment. See above.

>So get off my back

Absolutely not. You said something that was decidedly ignorant(how dare people praise x good thing done by omg horrible y people!), and I called you out on it. I expect better discussion and people skills from someone who holds position of a CTO rather than just "haha you're all paid shills!"

Re: Meta Llama 3

#958

Earlier quoted context omitted.

I just hosted both models here: https://chat.tune.app/ Playground: https://studio.tune.app/

Thanks for the link I just tested them and they also weark in europe without the need to start a VPN. What specs are needed to run these models. I mean the llama 70B and the Wizard 8Bx22 model. On your site they run very nicely and the answears they provide are really good they booth passed my small test and I would love to run one of them locally. So far I only ran 8B models on my 16GB RAM pc using LM Studio but hav…

Hey Christoph, thanks for trying it out - we're running this on the cloud, particularly GCP, on A100s (80g).

On your query about running these models locally, I'm not sure if just upgrading your RAM would have the same throughput as what you see on the website. You can upgrade your RAM but you might get pretty bad tokens/sec.

Re: Meta Llama 3

#959
post #4

They've got a console for it as well, https://www.meta.ai/ And announcing a lot of integration across the Meta product suite, https://about.fb.com/news/2024/04/meta-ai-assistant-built-wi... Neglected to include comparisons against GPT-4-Turbo or Claude Opus, so I guess it's far from being a frontier model. We'll see how it fares in the LLM Arena.

> Meta AI isn't available yet in your country Where is it available? I got this in Norway.

Anakin AI has Llama 3 models available right now: https://app.anakin.ai/

Re: Meta Llama 3

#960

Earlier quoted context omitted.

It's a very plausible rumor, but it is misleading in this context, because the rumor also states that it's a mixture of experts model with 8 experts, suggesting that most (perhaps as many as 7/8) of those weights are unused by any particular inference pass. That might suggest that GPT-4 should be thought of as something like a 250B model. But there's also some selection for the remaining 1/8 of weights that are used…

But from an output quality standpoint the total parameter count still seems more relevant. For example 8x7B Mixtral only executes 13B parameters per token, but it behaves comparable to 34B and 70B models, which tracks with its total size of ~45B parameters. You get some of the training and inference advantages of a 13B model, with the strength of a 45B model. Similarly, if GPT-4 is really 1.8T you would expect it to…

"For example 8x7B Mixtral only executes 13B parameters per token, but it behaves comparable to 34B and 70B models"

Are you sure about that? I'm pretty sure Miqu (the leaked Mistral 70b model) is generally thought to be smarter than Mixtral 8x7b.

Post reply on HN