Live data from Hacker News

DeepSeek V4 Pro beats GPT-5.5 Pro on precision

runtimewire.com

211–220 of 249 posts

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#211

It’s four poorly constructed arbitrary experiments which say very little about the competency of either model. The article reads like thin, auto-generated ai clickbait for nerd sniping or shilling a model. Consider the lead: > DeepSeek V4 Pro wins this head-to-head by being more exact where it matters: following instructions, matching schemas, and solving edge cases cleanly. GPT-5.5 Pro is still strong, but it gave a…

[dead]

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#212

I've been using deepseek v4 for cost/performance reasons. I feel it is generally not as good as some others, but in the end, you can make any model work by giving it the right acceptance criteria. Use detailed specs, use tests, and give it the power to iterate until it works. One-shot is a poor metric for performance.

I’m not sure all models will converge on your acceptance criteria. I’ve done quite a bit of varied agent based modeling and scientific modeling in that domain and just because you have some grounding to check against and some ideas on how you might go about getting to a convergence point doesn’t mean you’ll actually converge, you can absolutely get stuck in the information space iterating away, never finding your desired solutions.

It helps but you often have to step in the failure cases and guide them or forcibly fix certain paths to get a solution.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#213

Curious for folks who have made the switch I’m considering: if I swapped Claude Code to DeepSeek API pricing, would I get more bang for my buck compared to the $100 Max plan I’m using now? I only hit the 5 hour limit every few days and the weekly limit a day or two before it resets at the most aggressive. I wouldn’t expect my usage to increase dramatically, other than not being stopped by limits. I’m still apprehensi…

If you worry about sending your data off for inference, Fireworks is one of the companies serving open models with solid performance and compliance/zero data retention sorted out. OpenCode supports them and many others. Cursor uses them. They don't have the super-cheap cache reads deal that DeepSeek's own endpoint does, but are still well below Anthropic API rates. (Though crucially you're not paying API rates now!)

DeepSeek and Xiaomi's deals on cache reads go with their models' latest gens making caching cheaper (using less space for KVs). No open-model inference provider has decided to match the pricing. I'm sure that says something about how inference pricing works, but not completely sure what.

Agree with others that top open models aren't on the frontier, and I would expect differences doing big-picture planning or anywhere you're only giving broad brushstrokes and looking for a lot to be guessed. But they do seem fine at coding from a a concrete plan! No experience in huge codebases because I only use them outside work, but they seem good enough about gathering info before they dive in that I'd expect them to grep around as they need.

An annoying caveat: individual subscription plans, used heavily, are much cheaper than the API -- see https://she-llac.com/claude-limits -- which complicates any argument about cost. I still think open models are worth playing with. They're one of the things that let us treat this as a technology rather than just as the product offerings of one of a few companies.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#214

It’s four poorly constructed arbitrary experiments which say very little about the competency of either model. The article reads like thin, auto-generated ai clickbait for nerd sniping or shilling a model. Consider the lead: > DeepSeek V4 Pro wins this head-to-head by being more exact where it matters: following instructions, matching schemas, and solving edge cases cleanly. GPT-5.5 Pro is still strong, but it gave a…

> poorly constructed arbitrary experiments which say very little about the competency of either model. No one ever says this about the “pelican on a bicycle” metric

Actually, simonw has started saying that after qwen 27B beat Opus 4.7

https://news.ycombinator.com/item?id=48446348

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#215

Earlier quoted context omitted.

Yeah, the discounted deepseek inference is subsidized by the CCP for a reason, and it's one that might well come back to bite.

> deepseek inference is subsidized by the CCP What is that claim based on?

Besides common sense given the clear geopolitical context, sources like:

[1] https://chinaselectcommittee.house.gov/sites/evo-subsites/se... [2] https://ai.americansecurityproject.org/news/ai-imperative-20...

and more.

Of course, you can choose to ignore America-biased sources, but since it aligns with the obvious.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#216

Earlier quoted context omitted.

Yeah, the discounted deepseek inference is subsidized by the CCP for a reason, and it's one that might well come back to bite.

There is no evidence it is subsidized. Actually, there is evidence that (1) electricity is cheap in China & (2) deepseek is a very efficient model.

I think there is sufficient evidence to think its very likely. For example: https://www.americansecurityproject.org/wp-content/uploads/2...

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#217

Earlier quoted context omitted.

There are monied interests that do not want inexpensive Chinese successors to Scam Altman's creation.

They're inexpensive because they're derived from his creation.

The creation--which isn't "his" in the first place, by any standard definition--was not only itself "derived from" our creations but was always supposed to be "open".

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#218

Earlier quoted context omitted.

> deepseek inference is subsidized by the CCP What is that claim based on?

Besides common sense given the clear geopolitical context, sources like: [1] https://chinaselectcommittee.house.gov/sites/evo-subsites/se... [2] https://ai.americansecurityproject.org/news/ai-imperative-20... and more. Of course, you can choose to ignore America-biased sources, but since it aligns with the obvious.

There is no evidence in those sources that DeepSeek is "subsidized" by the CCP in the way people imply (e.g. in an actively malicious*, market-distorting way that undercuts the competition, early Uber-style). They do receive tax breaks for their R&D research, a very common scheme in Europe (and which also used to be the case in the US, I believe). They also have public-private partnerships, e.g. the state is one of their clients. Also common in every free market economy. (SpaceX anyone?)

*This does not invalidate other concerns (censorship, privacy) but the way people phrase it makes it look like DeepSeek and co. are 'cheating' somehow with their business model by 'distorting' inference cost to make it way artificially lower than its 'natural price' (either notion being hopelessly naive)

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#219

Earlier quoted context omitted.

I think you've misunderstood the purpose of a lead (sic). Per Merriam-Webster [^1], a lede is: > the introductory section of a news story that is intended to entice the reader to read the full story (Emphasis mine) You may prefer more matter-of-fact phrasing, of course, but criticising a lede for attempting to achieve its goal is unjustified. [^1]: https://www.merriam-webster.com/dictionary/lede

A 'lede' is just an intentionally differentiated spelling of 'lead'; the origin of the word is just lead . Collins dictionary defines lede: a variant spelling of lead

Is it not an intentional spelling in order to coin journalistic jargon?
Post reply on HN