Live data from Hacker News

DeepSeek V4 – almost on the frontier

simonwillison.net

351–360 of 420 posts

Re: DeepSeek V4 – almost on the frontier

#351

Earlier quoted context omitted.

Open models running locally is the answer. Relying on proprietary, closed software always puts that company's priorities above your own when using their software. You have given up control. While running them locally presently doesn't make sense economically, you don't need to run them locally to address this issue. There is a lot of competition in hosting open models and you have a variety of services to choose from…

It'll be a while yet before open models that're good enough will be viable for local use. Heck I've been trying to use the Qwen 3.5 39B A3B on my system, which is modest but no slouch, and have only been able to get ~4.5 tok/s after optimization, and it really runs my system red (fans instantly go crazy). It's just not practical for serious work.

I've been using Qwen 3.5 and then 3.6 27b Q4 on Ollama with a single 7900 XTX with the codex cli, and I have been blown away by how genuinely useful it is. I've been able to ask it to do long, multi step problems, and it's able to do things that would have likely taken me days to iron out in a matter of hours, or even minutes sometimes.

I get about 30 tok/s, which is far from blazing, but given the capability it has it is absolutely viable for accelerating my work.

Re: DeepSeek V4 – almost on the frontier

#352
post #273

Earlier quoted context omitted.

> It obviously went through lots of files in both prompts but total cost? Just $0.09 for the Pro version. When people say that LLMs aren't worth it, it kills me. A lot of us, on average, make $100+ an hour. $0.09 is You can't even read the vast majority of prompt responses that fast. LLMs will continue to get better (I'm doubtful at previous rates, all indications are showing that progress is slowing and costs are in…

I know I'm guilty of making this sort of argument sometimes, but it's just not valid. I don't get paid for every waking hour of every day. Often I'm using an LLM for something that's uncompensated, so my hourly wage equivalent is irrelevant. And for times when we might use an LLM for something related to paid work, it's still money out of your paycheck (unless the employer is paying for it; go nuts in that case). And…

I don't understand, your employer doesn't pay for your AI use? If my employer didn't pay for it I just wouldn't use it at all out of principle. Just as I don't buy my own work laptop

Re: DeepSeek V4 – almost on the frontier

#353
post #9

Deepseek v4 Pro feels like Claude Opus 4.6 in it's personality but here's what I did find out about costs: I did cut loose Deepseek v4 on a decent sized Typescript codebase and asked it to only focus on a single endpoint and go in depth on it layer by layer (API, DTOs, service, database models) and form a complete picture of types involved and introduced and ensure no adhoc types are being introduced. It developed a…

> It obviously went through lots of files in both prompts but total cost? Just $0.09 for the Pro version. When people say that LLMs aren't worth it, it kills me. A lot of us, on average, make $100+ an hour. $0.09 is You can't even read the vast majority of prompt responses that fast. LLMs will continue to get better (I'm doubtful at previous rates, all indications are showing that progress is slowing and costs are in…

Biggest issue with Opus for me is not so much that it's expensive (though it is), but the fact it's slow especially during US working hours.

I prefer using slightly worse but significantly quicker models on a tighter leash and iterating faster, feels more productive

Re: DeepSeek V4 – almost on the frontier

#354
post #138
post #113

The biggest differentiator for me: DeepSeek just does what I ask. I've tried using both GPT and Claude for reverse engineering recently, both refused. I even got a warning on my OpenAI account.

We have an enterprise cursor account so I can try all the mainstream models. Using composer 2 on our own code which I obviously have the source code for I couldn't get it to turn on a debug flag to bypass license checks while I was troubleshooting something. Infuriating. It was like that old Patrick from SpongeBob meme. I don't understand why we would turn the models into law enforcement officers. Things that are ill…

Software engineering is one thing but if you look 10-20 years into the future and everyone can run models equivalent to today's SoTA locally with zero monitoring or censorship, that could... not be good.

Some people will use them responsibly but a lot of people will not.

LLMs are already frying some people's brains and there are some human desires that should not be encouraged

Re: DeepSeek V4 – almost on the frontier

#355
post #113

The biggest differentiator for me: DeepSeek just does what I ask. I've tried using both GPT and Claude for reverse engineering recently, both refused. I even got a warning on my OpenAI account.

> I even got a warning on my OpenAI account. This is kind of terrifying to me, regularly. No real manner of recourse to normal people without a following, potential exclusion from real fundamental tooling. Imagine OpenAI goes on to buy 20 companies and now you cant use Figma, Next, whatever just because you once tripped some very foggy line somehow. Not just OpenAI but the entire ecosystem is so... hard to read. I wa…

I think it’s so bizarre that chatgpt regularly gives me advice on how to get around it’s filters. Like, literally “I can’t do anything if you use copyrighted character’s name, but how about you just say ‘someone that looks like character’”. If you are going to do that, can you just execute the instruction?

Re: DeepSeek V4 – almost on the frontier

#356

Earlier quoted context omitted.

"These safety rails" was referring to LLMs, which have far more nuanced and capable safety rails than chemical caps do, and accordingly also have much more assertive ways to enforce them.

It's the same underlying principle. If I want to ask a software tool what the suicide rate is for my county, I do not expect it to come back with: "Naughty boy! You said an unsafe word! You're getting a strike, and if you get two more, you're banned." This is totally out of the ordinary for a software product, and is absolutely a modern invention. Replace "suicide" with whatever the "AI Safety" obsession word is toda…

> If I want to ask a software tool what the suicide rate is for my county, I do not expect it to come back with: "Naughty boy! You said an unsafe word! You're getting a strike, and if you get two more, you're banned."

Did this happen?

I just tested this query in Grok, Gemini, Claude, and ChatGPT and 0% of them admonished me or refused to return an answer.

Just like every single conversation I've ever had on this topic, you have to make up examples that aren't even true. Why don't you just share what you were doing that you feel you were unfairly prevented from?

(I have an inkling why you won't do that...)

Re: DeepSeek V4 – almost on the frontier

#357

So RPI/QRSPI like skills (e.g. https://github.com/mattpocock/skills and https://github.com/humanlayer/humanlayer/tree/main/.claude/c... and https://github.com/dfrysinger/qrspi-plus ) for working with claude code work well enough for me that they can reliably* produce code that matches the plan/spec in a way they did not till December 2025. I have a gut feeling that these models can do just as well, has someone run a…

are you talking about a single prompt that runs for 24 hours or 8 hours of developer time spent in a single session?

Re: DeepSeek V4 – almost on the frontier

#358

Earlier quoted context omitted.

"the state" is just shorthand we use for "other people in my community" > I'll thank you and your kind for needing to distractedly tap the "Agree" button on my car's infotainment every time I start it to confirm that I will pay attention to the road. Does that actually mitigate antisocial usecases? No? Then it's not what I'm talking about :) Of course if you wanted to you could just share specifically what totally-re…

> "the state" is just shorthand we use for "other people in my community" It's a very different abstraction layer, in the same way as individual cells vs the entity that is you. The entity that comes together from all those "other people in my community" and its priorities are different to the individual desires. > Does that actually mitigate antisocial usecases? No? Then it's not what I'm talking about :) Maybe it d…

Yes if you apply some logic to such extremity that it produces bad outcomes then you should stop applying that logic to those extreme cases.

Re: DeepSeek V4 – almost on the frontier

#359
post #14

I'm surprised that people here don't care at all about these models openly training on your data, especially if you use them straight from the model developer. Whereas things like "GitHub now automatically opts everyone into using their code for model training" get hundreds of justifiably angry comments, I never see this brought up anymore on posts like these talking about using Chinese models through OpenRouter. Thi…

I never see the output of the Claude or MS models without having to pay for the privilege. All the Chinese models are open weight and open source.

Re: DeepSeek V4 – almost on the frontier

#360
post #73

Earlier quoted context omitted.

You definitely have a bone to pick. Chinese researchers usually have given the world the most cheap and consistent high quality research around LLMs. They don't pretend, they do the work and release the goodies. Mostly so cheap, every one in the world has a chance to use close to frontier models. Why would you respond with "Anger"? You let us know what your real complaint is about and let's not feign indignation at o…

You're making completely unfounded assumptions about me. I use Chinese models myself.

Anthropic and OpenAI took your data, trained their model, and tell you "we are not going to tell you anything how we trained our models, we are not giving your the weights our models, you will have to pay us to access the model trained from your data".

they took your rights and your data.

Chinese labs took your data, trained their model, and tell you "this paper details how our models are trained using your data, here is the final weights of our model trained from your data, feel free to use it for what you want, it is your model trained on your data".

they converted your data, everything is still in your hand under your control.

you couldn't see the difference?

Your specific question can actually be translated as -

1. why people don't stop Chinese labs so US monopoly can be maintained?

2. why people don't stop Chinese labs providing free models to those who would otherwise never be able to afford the same $200 USD/month Anthropic and OpenAI subscriptions.

3. why people don't complain Chinese labs publishing those trillion dollar secret ideas on model training.

well, because most people are not dickhead I guess?

Post reply on HN