Live data from Hacker News

DeepSeek v4

api-docs.deepseek.com

621–630 of 1001 posts

Re: DeepSeek v4

#621
post #612
post #597

Something is odd with this model, their blog posts shows REALLY good results, but in most other third-party benchmarks, people realize it's not really SOTA, even bellow Kimi K2.6 and GLM-5/5.1 In my tests too[0], it doesn't reach top 10. One issue, which they also mentioned in their post, is that they can't really serve well the model at the moment, so V4-Pro is heavily rate-limited and gives a lot of timeout errors…

Hmm, the Flash performs significantly better than Pro in the benchmark? That's very strange; could rate limiting cause that?

Yes, Flash doesn't seem to have the same rate limits as Pro.

I expect once the API issues are fixed, for v4-pro to be around the same level as GLM-5.

Re: DeepSeek v4

#622

There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…

Very interesting. I wonder how much of this is due to the context length. I am unclear on the implementation strategy, you ran this problem as a 1-shot using chat mode, or using each on an agent harness?

Re: DeepSeek v4

#623

Seriously, why can't huge companies like OpenAI and Google produce documentation that is half this good?? https://api-docs.deepseek.com/guides/thinking_mode No BS, just a concise description of exactly what I need to write my own agent.

Because they produce revenue from products which abstract this away

Re: DeepSeek v4

#624

Earlier quoted context omitted.

What do you run these on? I've gotten comfortable with Claude but if folks are getting Opus performance for cheaper I'll switch.

Try Charm Crush first, it's a native binary. If it's unbearable, try opencode, just with the knowledge your system will probably be pwned soon since it's JS + NPM + vibe coding + some of the most insufferable devs in the industry behind that product. If you're feeling frisky, Zed has a decent agent harness and a very good editor.

I've downloaded Zed but haven't used it much, maybe this is my sign. Thanks!

Re: DeepSeek v4

#625

Earlier quoted context omitted.

On-device is incredibly far away from being viable. A $20 ChatGPT subscription beats the hell out of the 8B model that a $1,000 computer can run. Nvidia's forward PE ratio is only 20 for 2026. That's much lower than companies like Walmart and Costco. It's also growing nearly 100% YoY and has a $1 trillion backlog. I think Nvidia is cheap.

This is an assessment of the moment. When rate of AI data center construction slows down, then P/E will start to grow. Or are we saying that the pace will only grow forever? There are already signs of a slowdown in construction.

What are these signs you are referencing? Source?

Re: DeepSeek v4

#626

Earlier quoted context omitted.

I sometimes wonder if there are any security risks with using Chinese LLMs. Is there?

Backdooring software at scale. Spearphishing. Building reliance and exploiting it, through state subsidies, dumping, and market manipulation. Handicapping provision to the west for competitive advantage.

Anyone can do that via the scrapers. The model developers actually have something to lose tho

Re: DeepSeek v4

#627
post #540

I like deepseek. It works very well. I haven't tried v4 yet but on their web chat interface, just typing "Taiwan" causes it to give you a lecture about how Taiwan is part of China. :)

It's open source, so just delete those parameters. /s

Re: DeepSeek v4

#628

Assuming it is almost as good as Opus 4.6 (which benchmarks seem to give evidence for), and assuming we are having a good enough harness (PI, OpenCode), it's is now more than 5x cheaper. I just want to remind you that this is happening at the same time as Anthropic A/B tests removal of Code from Pro Plan, and as OpenAI releases gpt-5.5 2x more expensive than gpt-5.4...

> Assuming it is almost as good as Opus 4.6 (which benchmarks seem to give evidence for) That’s a big if. It’s my experience that models that perform very well on benchmarks do not necessarily perform well in real life. I’ve mostly started ignoring the benchmarks and run my own evals.

If benchmarks are all to be believed then gemini 3.1 and grok 4.2 are still in the lead pack. A laughable notion to anyone who has actually tried to use them and compared.

Re: DeepSeek v4

#629
post #325

Earlier quoted context omitted.

I wonder why there aren't more open weights model with support for prompt caching on OpenRouter.

It is tricky to build good infrastructure for prompt caching.

Its as simple as telling your claude code to implement prompt caching!

Re: DeepSeek v4

#630
post #102

Earlier quoted context omitted.

I didn't express this well but my interest isn't "who is in the top spot", and is more _why and _how various labs get the results they do. This is also magnified by the fact that I'm not only interested in hosted providers of inference but local models as well. What's your take on the best model to run for coding on 24GB of VRAM locally after the last few weeks of releases? Which harness do you prefer? What quants do…

Follow the AI newsletters. They bundle the news along with their Op-Ed and summarize it better.

[dead]
Post reply on HN