Live data from Hacker News

DeepSeek V4 – almost on the frontier

simonwillison.net

261–270 of 420 posts

Re: DeepSeek V4 – almost on the frontier

#261
post #8

I tweeted about some implementation and review runs that used V4 Pro. Even without the currently discounted pricing, the value is incredible. It takes about twice as long to finish code reviews given an identical context compared to opus 4.7/gpt 5.5 but at 1/10 the cost of less, there's just no comparison. https://twitter.com/aljosa/status/2049176528638902555

Did you do this test through OpenRouter?

Yes, but locked to the official DeepSeek provider since it's the only one that has the discounted pricing.

Re: DeepSeek V4 – almost on the frontier

#262
post #9

Deepseek v4 Pro feels like Claude Opus 4.6 in it's personality but here's what I did find out about costs: I did cut loose Deepseek v4 on a decent sized Typescript codebase and asked it to only focus on a single endpoint and go in depth on it layer by layer (API, DTOs, service, database models) and form a complete picture of types involved and introduced and ensure no adhoc types are being introduced. It developed a…

> It obviously went through lots of files in both prompts but total cost? Just $0.09 for the Pro version.

When people say that LLMs aren't worth it, it kills me.

A lot of us, on average, make $100+ an hour. $0.09 is You can't even read the vast majority of prompt responses that fast.

LLMs will continue to get better (I'm doubtful at previous rates, all indications are showing that progress is slowing and costs are increasing disproportionately).

It seems like >50% of devs think LLMs provide less than 0 value. I just do not get it.

Did they use an LLM one time 3 years ago and decide it's never going to be worth it? Have they even tried? Or have you only ever tried it on 1 giant, monolythic proprietary codebase where they're a total expert and decided that an LLM isn't as good as them, so it's "completely worthless"?

They are shockingly unhelpful on my company's codebase.

But that doesn't mean they are flat-out worthless.

Re: DeepSeek V4 – almost on the frontier

#263
I've found this to be a very good model, and I think I'd even go as far as rating it higher than Chatgpt.

ChatGPT has really degraded in my eyes, and I find Grok and Deepseek more helpful most of the time.

Of course, ChatGPT is better sometimes.

These models are just better than others at different cases, thus the reason to experiment.

Re: DeepSeek V4 – almost on the frontier

#264

Earlier quoted context omitted.

> I even got a warning on my OpenAI account. This is kind of terrifying to me, regularly. No real manner of recourse to normal people without a following, potential exclusion from real fundamental tooling. Imagine OpenAI goes on to buy 20 companies and now you cant use Figma, Next, whatever just because you once tripped some very foggy line somehow. Not just OpenAI but the entire ecosystem is so... hard to read. I wa…

It's probably because you were talking about a quote from a book (ie copyrighted material). Authors have sued the AI companies for repeating / memorizing copyrighted works, and getting an AI to discuss a quote would be making it repeat a portion of copyrighted work. Funny that your case is Kurt Vonnegut. I think I had Claude refuse a task where I was doing an OCR scan of a book review (in a zine / journal a family me…

Joseph Heller methinks, but probably not too far away in embedding space!

Re: DeepSeek V4 – almost on the frontier

#265

The pelican is really getting old as an a standalone evaluation metric. By now they are certainly going to be in training set if not explicitly tuned to produce it for the press on HN alone. Keep the pelican but isn’t it time to add something else more novel that all current and past models struggle with?

One shot canvas and svg images or animations are also just something that at this scale shouldn't be an issue at all, even Qwen running locally on 24gb cards can do impressive ones. Don't understand why this test gets any attention, I mean other than the pelicans which isn't a good test, theres no meat in this article.

And yet, look at the French one. Can't compete with one year old open weight models even though they just released a new model this week.

Re: DeepSeek V4 – almost on the frontier

#267

Earlier quoted context omitted.

> I don't understand why we would turn the models into law enforcement officers It's a simple corporate risk minimization strategy. Just look at how universally despised Grok is on HN. Not because it's a bad model, but because it has less aggressive alignment which means it can be coaxed into saying things that get Xai pilloried here and elsewhere.

It's mostly just a bad model. Plenty of people would be willing to overlook the baggage if the model was even marginally better than the competition.

I also used to see Grok boosting/slack-cutting on here/Reddit constantly back in Peak Subsidy when xAI was giving out hundreds of dollars of credits for free per month.

After they killed that and then stopped handing out free model access to users of every Cline fork for weeks following model releases, vibe coder hype moved back to Chinese models for cost and the SOTA models for quality.

Re: DeepSeek V4 – almost on the frontier

#268
post #82
post #47

Earlier quoted context omitted.

Anthropic's and OpenAI's costs seem to include a fairly ok margin, from the very fourth hand info I have.

In total, how many hands do you have?

If I was a betting man I'd bet that at least one of those hands is an LLM

Re: DeepSeek V4 – almost on the frontier

#269

Earlier quoted context omitted.

> I don't understand why we would turn the models into law enforcement officers It's a simple corporate risk minimization strategy. Just look at how universally despised Grok is on HN. Not because it's a bad model, but because it has less aggressive alignment which means it can be coaxed into saying things that get Xai pilloried here and elsewhere.

It's mostly just a bad model. Plenty of people would be willing to overlook the baggage if the model was even marginally better than the competition.

Agreed. There's are plenty of instances where people here on HN do mental gymnastics to justify using a truly good product when the company that builds it is morally bankrupt.

Not a criticism (I probably engage in that sort of thinking myself sometimes), just something I've observed. If Grok were actually good, we'd see that phenomenon here, but we don't.

Re: DeepSeek V4 – almost on the frontier

#270

Earlier quoted context omitted.

> All sorts of tools try to prevent dangerous/destructive uses But they don't threaten their users or have an "N strikes and you're out" policy. I take those safety caps off of all the chemicals in my garage because I'm a grown-ass adult and those caps are a pain in the butt. I would not expect the manufacturer of a solvent to show up at my house lecturing me about safety and threatening to ban me from buying his pro…

Sure but they would if they could. If they knew idiots were doing idiot things with their products (or evils doing evil things) and did not utilize available methods to prevent them, then the company ends up holding liability. And no, this is not easily signed away in a contract.

There actually is a very important distinction between "would if they could" and "they can and do", though.
Post reply on HN