Earlier quoted context omitted.
And they even allow you to use it without logging in. Didnt expect that from Meta.
Not in the EU though
Meta Llama 3
211–220 of 965 posts
Re: Meta Llama 3
#212Earlier quoted context omitted.
For the still training 400B: Llama 3 GPT 4(Published) BBH 85.3 83.1 MMLU 86.1 86.4 DROP 83.5 80.9 GSM8K 94.1 92.0 MATH 57.8 52.9 HumEv 84.1 74.4 Although it should be noted that the API numbers were generally better than published numbers for GPT4. [1]: https://deepmind.google/technologies/gemini/
Hm, how much VRAM would this take to run?
Re: Meta Llama 3
#213I downloaded llama3:8b-instruct-q4_0 in ollama and said "hi" and it answered with 10 screen long rant. This is an exerpt. > You're welcome! It was a pleasure chatting with you. Bye for now!assistant > Bye for now!assistant > Bye!assistant
Re: Meta Llama 3
#214How does it make monetary sense to release open source models? AFAIK it's very expensive to train them. Do Meta/Mistral have any plans to monetize them?
Before Llama, Meta was defined in the short-term by dubious investment in "metaverse" and cryptocurrency nonsense.
Now they are an open AI champion.
Re: Meta Llama 3
#215We've got an API out here: https://replicate.com/blog/run-llama-3-with-an-api You can also chat with it here: https://llama3.replicate.dev/
Re: Meta Llama 3
#216Earlier quoted context omitted.
Comparing to the numbers here https://www.anthropic.com/news/claude-3-family the ones of Llama 400B seem slightly lower, but of course it's just a checkpoint that they benchmarked and they are still training further.
Indeed. But if GPT-4 is actually 1.76T as rumored, an open-weight 400B is quite the achievement even if it's only just competitive.
Re: Meta Llama 3
#217Cloudflare AI team, any chance it’ll be on Workers AI soon? I’m sure some of you are lurking :)
Re: Meta Llama 3
#218Awesome, but I am surprised by the constrained context window as it balloons everywhere else. Am I missing something? 8k seems quite low in current landscape.
Honestly, I swear to god, been working 12 hours a day with these for a year now, llama.cpp, Claude, OpenAI, Mistral, Gemini: The long context window isn't worth much and is currently creating more problems than it's worth for the bigs, with their "unlimited" use pricing models. Let's take Claude 3's web UI as an example. We build it, and go the obvious route: we simply use as much of the context as possible, given ch…
Re: Meta Llama 3
#219What sort of hardware is needed to run either of these models in a usable fashion? I suppose the bigger 70B model is completely unusable for regular mortals...
Re: Meta Llama 3
#220Just got uploaded to HuggingFace: https://huggingface.co/meta-llama/Meta-Llama-3-8B https://huggingface.co/meta-llama/Meta-Llama-3-70B
Playground: https://studio.tune.app/