Phi-3 Technical Report
arxiv.org
Phi-3 Technical Report
1–10 of 132 posts
Re: Phi-3 Technical Report
#2And on LMSYS English, Llama 3 8B is on par with GPT-4 (not GPT-4-Turbo), as well as Mistral-Large.
Source: https://chat.lmsys.org/?leaderboard (select English in the dropdown)
So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones? Kinda? Wild.
(I'm sure there's a lot of nuance to it, for one these benchmarks are not so hard to game, we'll see how the dust settles, but still...)
Phi-3-mini 3.8b: 71.2
Phi-3-small 7b: 74.9
Phi-3-medium 14b: 78.2
Phi-2 2.7b: 58.8
Mistral 7b: 61.0
Gemma 7b: 62.0
Llama-3-In 8b: 68.0
Mixtral 8x7b: 69.9
GPT-3.5 1106: 75.3
(these are averages across all tasks for each model, but looking at individual scores shows a similar picture)
Re: Phi-3 Technical Report
#3Re: Phi-3 Technical Report
#4Incredible, rivals Llama 3 8B with 3.8B parameters after less than a week of release. And on LMSYS English, Llama 3 8B is on par with GPT-4 (not GPT-4-Turbo), as well as Mistral-Large. Source: https://chat.lmsys.org/?leaderboard (select English in the dropdown) So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones? Kinda? Wild. (I'm sure there's a lot of nuance to it, f…
Feels incredible to be living in a time with such neck breaking innovations. What are chances we’ll have a <100B parameter GPT4/Claude Opus model in the next 5 years?
Re: Phi-3 Technical Report
#5Incredible, rivals Llama 3 8B with 3.8B parameters after less than a week of release. And on LMSYS English, Llama 3 8B is on par with GPT-4 (not GPT-4-Turbo), as well as Mistral-Large. Source: https://chat.lmsys.org/?leaderboard (select English in the dropdown) So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones? Kinda? Wild. (I'm sure there's a lot of nuance to it, f…
Source?
Re: Phi-3 Technical Report
#6Incredible, rivals Llama 3 8B with 3.8B parameters after less than a week of release. And on LMSYS English, Llama 3 8B is on par with GPT-4 (not GPT-4-Turbo), as well as Mistral-Large. Source: https://chat.lmsys.org/?leaderboard (select English in the dropdown) So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones? Kinda? Wild. (I'm sure there's a lot of nuance to it, f…
>And on LMSYS English, Llama 3 8B is well above GPT-4 Source?
Re: Phi-3 Technical Report
#7Incredible, rivals Llama 3 8B with 3.8B parameters after less than a week of release. And on LMSYS English, Llama 3 8B is on par with GPT-4 (not GPT-4-Turbo), as well as Mistral-Large. Source: https://chat.lmsys.org/?leaderboard (select English in the dropdown) So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones? Kinda? Wild. (I'm sure there's a lot of nuance to it, f…
Per the paper, phi3-mini (which is english-only) quantised to 4bit uses 1.8gb RAM and outputs 1212 tokens/sec (correction: 12 tokens/sec) on iOS.
A model on par with GPT-3.5 running on phones!
(weights haven't been released, though)
Re: Phi-3 Technical Report
#8Earlier quoted context omitted.
>And on LMSYS English, Llama 3 8B is well above GPT-4 Source?
Right thanks for the reminder, I added it
I also don't think they "beat Llama 3 8B"; their own abstract says "rivals that of models such as Mixtral 8x7B and GPT-3.5", "rivals" not even "beats".
Great model, but let's not overplay it.