Live data from Hacker News

Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

techcommunity.microsoft.com

71–80 of 148 posts

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#71
I'm not too excited by Phi-4 benchmark results - It is#BenchmarkInflation.

Microsoft Research just dropped Phi-4 14B, an open-source model that’s turning heads. It claims to rival Llama 3.3 70B with a fraction of the parameters — 5x fewer, to be exact.

What’s the secret? Synthetic data. -> Higher quality, Less misinformation, More diversity

But the Phi models always have great benchmark scores, but they always disappoint me in real-world use cases.

Phi series is famous for to be trained on benchmarks.

I tried again with the hashtag#phi4 through Ollama - but its not satisfactory.

To me, at the moment - IFEval is the most important llm benchmark.

But look the smart business strategy of Microsoft:

have unlimited access to gpt-4 the input prompt it to generate 30B tokens train a 1B parameter model call it phi-1 show benchmarks beating models 10x the size never release the data never detail how to generate the data( this time they told in very high level) claim victory over small models

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#72
post #9

The most interesting thing about this is the way it was trained using synthetic data, which is described in quite a bit of detail in the technical report: https://arxiv.org/abs/2412.08905 Microsoft haven't officially released the weights yet but there are unofficial GGUFs up on Hugging Face already. I tried this one: https://huggingface.co/matteogeniaccio/phi-4/tree/main I got it working with my LLM tool like this: l…

> More of my notes on Phi-4 here: https://simonwillison.net/2024/Dec/15/phi-4-technical-report...

Nice. Thanks.

Do you think sampling the stack traces of millions of machines is a good dataset for improving code performance? Maybe sample android/jvm bytecode.

Maybe a sort of novelty sampling to avoid re-sampling hot-path?

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#73
post #62

Earlier quoted context omitted.

We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…

>MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. If only those models supported anything other than English

Is subtext to this uncensored Chinese support?

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#74

Earlier quoted context omitted.

We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…

[dead]

>> Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4

> I'm skeptical; the llama 3.1 405B model is the only comparable model I've used, and it's significantly larger than the 70B models you can run locally.

Every new Llama generation achieved to beat larger models of the previous generation with smaller ones.

Check Kagi's LLM benchmark: https://help.kagi.com/kagi/ai/llm-benchmark.html

Check the HN thread around the 3.3 70b release: https://news.ycombinator.com/item?id=42341388

And their own benchmark results in their model card: https://github.com/meta-llama/llama-models/blob/main/models%...

Groq's post about it: https://groq.com/a-new-scaling-paradigm-metas-llama-3-3-70b-...

Etc

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#75

Earlier quoted context omitted.

How many parameters did ChatGPT have in Dec 2022 when it first broke into mainstream news?

GPT-3 had 175B, and the original ChatGPT was probably just a GPT-3 finetune (although they called it gpt-3.5, so it could have been different). However, it was severely undertrained. Llama-3.1-8B is better in most ways than the original ChatGPT; a well-trained ~70B usually feels GPT-4-level. The latest Llama release, llama-3.3-70b, goes toe-to-toe even with much larger models (albeit is bad at coding, like all Llama…

> However, it was severely undertrained

by modern standards. at the time, it was trained according to neural scaling laws oai believed to hold.

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#76

Earlier quoted context omitted.

We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…

[dead]

> Is a $8000 MBP regular consumer hardware? If you don't think so, then the answer is probably no.

The very first Apple McIntosh was not far from that price at its release. Adjusted for inflation of course.

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#77

Earlier quoted context omitted.

We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…

The coolness of local LLMs is THE only reason I am sadly eyeing upgrading from M1 64GB to M4/5 128+GB.

I bought an old used desktop computer, a used 3090, and upgraded the power supply, all for around 900€. Didn't assemble it all yet. But it will be able to comfortably run 30B parameter models with 30-40 T/s. The M4 Max can do ~10 T/s, which is not great once you really want to rely on it for your productivity.

Yes, it is not "local" as I will have to use the internet when not at home. But it will also not drain the battery very quickly when using it, which I suspect would happen to a Macbook Pro running such models. Also 70B models are out of reach of my setup, but I think they are painfully slow on Mac hardware.

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#78

For prompt adherence it still fails on tasks that Gemma2 27b nails every time. I haven't been impressed with any of the Phi family of models. The large context is very nice, though Gemma2 plays very well with self-extend.

It's a much smaller model though.

I think the point is more the demonstration that such a small model can have such good performance than any actual usefulness.

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#79

Earlier quoted context omitted.

We’re already past that point! MacBooks can easily run models exceeding GPT-3.5, such as Llama 3.1 8B, Qwen 2.5 8B, or Gemma 2 9B. These models run at very comfortable speeds on Apple Silicon. And they are distinctly more capable and less prone to hallucination than GPT-3.5 was. Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I ha…

The coolness of local LLMs is THE only reason I am sadly eyeing upgrading from M1 64GB to M4/5 128+GB.

I'm returning my 96GB m2 max. It can run unquantized llama 3.3 70B but tokens per second is slow as molasses and still I couldn't find any use for it, just kept going back to perplexity when I actually needed to find an answer to something.

Re: Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning

#80
post #47
post #39

Earlier quoted context omitted.

Appreciate your rapid analysis of new models, Simon. Have any models you've tested performed well on the pelican SVG task?

gemini-exp-1206 is my new favorite: https://simonwillison.net/2024/Dec/6/gemini-exp-1206/ Claude 3.5 Sonnet is in second place: https://github.com/simonw/pelican-bicycle?tab=readme-ov-file...

They probably trained it for this specific task (generating SVG images), right?
Post reply on HN