Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
61–70 of 148 posts
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#62Looks like it punches way above its weight(s). How far are we from running a GPT-3/GPT-4 level LLM on regular consumer hardware, like a MacBook Pro?
We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…
If only those models supported anything other than English
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#63The most interesting thing about this is the way it was trained using synthetic data, which is described in quite a bit of detail in the technical report: https://arxiv.org/abs/2412.08905 Microsoft haven't officially released the weights yet but there are unofficial GGUFs up on Hugging Face already. I tried this one: https://huggingface.co/matteogeniaccio/phi-4/tree/main I got it working with my LLM tool like this: l…
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#64Earlier quoted context omitted.
We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…
>MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. If only those models supported anything other than English
All of the Qwen models are basically fluent in both English and Chinese.
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#65Looks like it punches way above its weight(s). How far are we from running a GPT-3/GPT-4 level LLM on regular consumer hardware, like a MacBook Pro?
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#66Earlier quoted context omitted.
>Microsoft haven't officially released the weights Thought it was official just not on huggingface but rather whatever azure competitor thing they're pushing?
I found their AI Foundry thing so hard to figure out I couldn't tell if they had released weights (as opposed to a way of running it via an API). Since there are GGUFs now so someone must have released some weights somewhere.
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#67Earlier quoted context omitted.
>MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. If only those models supported anything other than English
Llama 3.1 8B advertises itself as multilingual. All of the Qwen models are basically fluent in both English and Chinese.
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#68Earlier quoted context omitted.
We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…
[dead]
Quite impressive.
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#69Earlier quoted context omitted.
I've seen people using "draw a unicorn using tikz" https://adamkdean.co.uk/posts/gpt-unicorn-a-daily-exploratio...
That came, IIRC, from one of the OpenAI or Microsoft people (Sebastian Bubeck); it was recounted in an NPR podcast "Greetings from Earth" https://www.thisamericanlife.org/803/transcript
The most significant part I took away is that when safety "alignment" was done the ability plummeted. So that really makes me wonder how much better these models would be if they weren't lobotomized to prevent them from saying bad words.
Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
#70Earlier quoted context omitted.
We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…
>MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. If only those models supported anything other than English