Live data from Hacker News

DeepSeek V4 – almost on the frontier

simonwillison.net

311–320 of 420 posts

Re: DeepSeek V4 – almost on the frontier

#311
post #270

Earlier quoted context omitted.

There actually is a very important distinction between "would if they could" and "they can and do", though.

Uhh right, but describing that as "dystopian" is frankly hysterical. It's an obvious corollary of good things (like product liability). Virtually everyone I've heard complain about these safety rails was up to antisocial (at best) stuff. I've never heard a sympathetic use-case. It's objectively good that companies can be held responsible for misuse of their products and that they are therefore incentivized to mitigat…

I don't think that "dystopian" necessarily goes far enough, this would be one of the rare times where I would call it a fascist mentality - the idea that everything's primary allegiance is to the state and the goals of the state rather than those of the customer or the user.

I want a default that has people empowered, rather than something where it's just another performative smokescreen caused by overzealous product liability. I'll thank you and your kind for needing to distractedly tap the "Agree" button on my car's infotainment every time I start it to confirm that I will pay attention to the road.

Re: DeepSeek V4 – almost on the frontier

#312

Earlier quoted context omitted.

Well, the context was running the models via open router, not hosting 800B> models yourself. Of course, if given the option I believe most people would pick ”don’t share sensitive data”. What I’m trying to say is that EVERYONE uses your data, even the sensitive type. So you might aswell use an endpoint that does what it says and treat EVERY endpoint whether that’s OpenAI or anthropic as if it’s collecting all of your…

No, not everyone uses your data. There are providers who very explicitly do not collect or use your data.

Sure, and I won’t collect or otherwise store your credit card info if you send it to me. Trust me bro :)

No but seriously, I am astonished by the level of trust you have for these for-profit companies. I’ll remind you of this quote:

”Zuckerberg: People just submitted it. Zuckerberg: I don't know why. Zuckerberg: They "trust me" Zuckerberg: Dumb fucks”

Re: DeepSeek V4 – almost on the frontier

#313
post #138
post #113

The biggest differentiator for me: DeepSeek just does what I ask. I've tried using both GPT and Claude for reverse engineering recently, both refused. I even got a warning on my OpenAI account.

We have an enterprise cursor account so I can try all the mainstream models. Using composer 2 on our own code which I obviously have the source code for I couldn't get it to turn on a debug flag to bypass license checks while I was troubleshooting something. Infuriating. It was like that old Patrick from SpongeBob meme. I don't understand why we would turn the models into law enforcement officers. Things that are ill…

> Things that are illegal are still illegal and we have professionals to deal with crimes.

This is quite naive take though. The direction of travel is more fascism in Western governments where duties of traditional policing are taken over by big corporations whilst police forces are being gutted and made impotent.

Re: DeepSeek V4 – almost on the frontier

#314

I tried DeepSeek via chat, and gave it a rather simple question: "Can you tell me who was on series 8 of Taskmaster, and what's the general opinion about the series? No spoilers!" It told me amongst other things that Paul Sinha was diagnosed with Parkinsons, as well as who the winner was. Then I said, "But I said no spoilers!" And it apologised for telling me Paul Sinha was diagnosed with Parkinsons.

I was not able to reproduce your problem with that prompt, but I might have a reason for why you got that answer.

Did you enable reasoning ("DeepThink")? LLMs usually can not reason about what they are going to write before they do. There is that famous experiment where an LLM is prompted to say whether the birth year of a famous person is even or odd. If the LLM is constrained to only answer with "even" or "odd", the accuracy is around 50%, i.e. no better than random chance, but if the LLM is allowed to first answer with the birth year of the famous person followed by whether the year is even or odd, it is able to "see" what the year is, and answers correctly almost every time.

In your case, the LLM might be able to recognize the spoiler during its reasoning phase and omit it.

Another explanation might be that the LLM interpreted the "No spoilers!" as "Do not spoil the tasks of the show" instead of "Do not spoil the winner".

Lastly, the question "Can you tell me...?" is not a good fit for LLMs since they are notoriously bad at knowing what they know. You can leave it out to save a few characters.

Re: DeepSeek V4 – almost on the frontier

#315

Earlier quoted context omitted.

Your method of combining models to strengthen the implementation reminds me of how we form stronger alloys by combining metals!

it also sounds like a lot to manage, do you have some sort of agentic framework that's treating all of these llm's you have access to as sort of inputs that it optimizes?

Not op but I wrote llm-consortium to prompt multiple models and create a synthesis. And it can run on an openai endpoint using llm-model-gateway. It's expensive, naturally, but for situations where you absolutely must get max intelligence its hard to beat.

e.g.

  Pelican Riding a Bicycle — Engineering Study by DeepSeek v4 Pro, Kimi K2.6, and GLM-5.1 (1 iteration in synthesis mode with DeepSeek v4 flash as judge)
https://htmlpreview.github.io/?https://gist.githubuserconten...

Re: DeepSeek V4 – almost on the frontier

#316
post #113

The biggest differentiator for me: DeepSeek just does what I ask. I've tried using both GPT and Claude for reverse engineering recently, both refused. I even got a warning on my OpenAI account.

This is huge for me too, I was working on something super benign the other day and GPT flagged it for Cyber risk, Deepseek just does the work, its fast and cheap. Its only missing image support IMO, once deepseek cracks image too its going to be hard for anthropic and openai to compete.

Re: DeepSeek V4 – almost on the frontier

#317

Earlier quoted context omitted.

> I even got a warning on my OpenAI account. This is kind of terrifying to me, regularly. No real manner of recourse to normal people without a following, potential exclusion from real fundamental tooling. Imagine OpenAI goes on to buy 20 companies and now you cant use Figma, Next, whatever just because you once tripped some very foggy line somehow. Not just OpenAI but the entire ecosystem is so... hard to read. I wa…

Open models running locally is the answer. Relying on proprietary, closed software always puts that company's priorities above your own when using their software. You have given up control. While running them locally presently doesn't make sense economically, you don't need to run them locally to address this issue. There is a lot of competition in hosting open models and you have a variety of services to choose from…

You don't need to run the model locally if you don't care about sharing your data. Personally I am happy to share data with Kimi or Deepseek if it means we get better OSS models. For private stuff though local is king

Re: DeepSeek V4 – almost on the frontier

#318

Earlier quoted context omitted.

Uhh right, but describing that as "dystopian" is frankly hysterical. It's an obvious corollary of good things (like product liability). Virtually everyone I've heard complain about these safety rails was up to antisocial (at best) stuff. I've never heard a sympathetic use-case. It's objectively good that companies can be held responsible for misuse of their products and that they are therefore incentivized to mitigat…

I don't think that "dystopian" necessarily goes far enough, this would be one of the rare times where I would call it a fascist mentality - the idea that everything's primary allegiance is to the state and the goals of the state rather than those of the customer or the user. I want a default that has people empowered, rather than something where it's just another performative smokescreen caused by overzealous product…

"the state" is just shorthand we use for "other people in my community"

> I'll thank you and your kind for needing to distractedly tap the "Agree" button on my car's infotainment every time I start it to confirm that I will pay attention to the road.

Does that actually mitigate antisocial usecases? No? Then it's not what I'm talking about :)

Of course if you wanted to you could just share specifically what totally-reasonable LLM use-case you have in mind that's neutered by this "fascist mentality" instead of dreaming up unrelated instances.

Re: DeepSeek V4 – almost on the frontier

#320
post #113

The biggest differentiator for me: DeepSeek just does what I ask. I've tried using both GPT and Claude for reverse engineering recently, both refused. I even got a warning on my OpenAI account.

Claude has refused to run nmap so I can locate my own computer on my own network! The guard rails are completely out of control.
Post reply on HN