Live data from Hacker News

DeepSeek 4 Flash local inference engine for Metal

github.com

11–20 of 171 posts

Re: DeepSeek 4 Flash local inference engine for Metal

#11
post #6

Earlier quoted context omitted.

There will always be a huge gap between frontier models and open source models (unless you're very rich). This whole industry makes no sense, everyone is ignoring the unit economics. It cost 20k a month to running Kimi 2.6 at decent tok/ps, to sell those tokens at a profit you'd need your hardware costs to be less 1k a month. Everyone who's betting their competency on the generosity of billionaires selling tokens for…

If you looked at a graph of GPU power in consumer hardware and model capability per billion parameters over time, it seems inevitable that in the next few years a "good enough" model will run on entry-level hardware. Of course there will always be larger flagship models, but if you can count on decent on-device inference, it materially changes what you can build.

[flagged]

Re: DeepSeek 4 Flash local inference engine for Metal

#12
post #11

Earlier quoted context omitted.

If you looked at a graph of GPU power in consumer hardware and model capability per billion parameters over time, it seems inevitable that in the next few years a "good enough" model will run on entry-level hardware. Of course there will always be larger flagship models, but if you can count on decent on-device inference, it materially changes what you can build.

[flagged]

No offense, this is a crazy worthless contribution to the discussion.

Why?

Re: DeepSeek 4 Flash local inference engine for Metal

#13
I am curious about it producing less tokens except for the max mode. I love DeepSeek V4 Flash and I use it extensively, it's so cheap I can use it all day and still not use all my 10$ OpenCode Go subscription. I use it always in max mode because of this, but now I wonder whether I should rather use high.

Re: DeepSeek 4 Flash local inference engine for Metal

#14
post #6
post #3

This is so sick. I'm really curious to see what focused effort on optimizing a single open source model can look like over many months. Not only on the inference serving side, but also on the harness optimization side and building custom workflows to narrow the gap between things frontier models can infer and deduce and what open source models natively lack due to size, training etc.

There will always be a huge gap between frontier models and open source models (unless you're very rich). This whole industry makes no sense, everyone is ignoring the unit economics. It cost 20k a month to running Kimi 2.6 at decent tok/ps, to sell those tokens at a profit you'd need your hardware costs to be less 1k a month. Everyone who's betting their competency on the generosity of billionaires selling tokens for…

Most tasks do not require frontier models, so as long as these models cover 95-99 per cent of the tasks, closed frontier models can be left for niche and specialized cases that are harder.

Re: DeepSeek 4 Flash local inference engine for Metal

#17
post #15

A random, funny, interesting and telling data point: my MacBook M3 Max while DS4 is generating tokens at full speed peaks 50W of energy usage...

"Data centers for LLMs are technically more energy efficient per-user than self-hosting LLM models due to economies-of-scale" is a data point the internet isn't ready for.

Re: DeepSeek 4 Flash local inference engine for Metal

#18
post #15

A random, funny, interesting and telling data point: my MacBook M3 Max while DS4 is generating tokens at full speed peaks 50W of energy usage...

I think I’ve seen about 60 watt total system whenever I’ve used a local model on a MacBook Pro or a Mac Studio. Baseline for the Mac Studio is like 10 W and like 6 W for the MacBook Pro.

Re: DeepSeek 4 Flash local inference engine for Metal

#19
So just gonna ask a question, probably will get downvoted

I know this is flash, but….

But other than this guy, did our whole society seriously never flamegraph this stuff before we started requesting nuclear reactors colocated at data centers and like more than 10% of gdp?

Someone needs to answer because this isn’t even a m4 or m5… WHAT THE FUCK

Sidenote: shout out antirez love my redis :)

Re: DeepSeek 4 Flash local inference engine for Metal

#20

So just gonna ask a question, probably will get downvoted I know this is flash, but…. But other than this guy, did our whole society seriously never flamegraph this stuff before we started requesting nuclear reactors colocated at data centers and like more than 10% of gdp? Someone needs to answer because this isn’t even a m4 or m5… WHAT THE FUCK Sidenote: shout out antirez love my redis :)

DSv4 generates much faster on NVIDIA class hardware. It is just a very efficient model.
Post reply on HN