Live data from Hacker News

DeepSeek v4

api-docs.deepseek.com

251–260 of 1001 posts

Re: DeepSeek v4

#251
Feels like the real story here is cost/performance tradeoff rather than raw capability. Benchmarks keep moving incrementally, but efficiency gains like this actually change who can afford to build on top.

Re: DeepSeek v4

#252

Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…

But remember to not ask about Taiwan!

Re: DeepSeek v4

#253
post #96

MMLU-Pro: Gemini-3.1-Pro at 91.0 Opus-4.6 at 89.1 GPT-5.4, Kimi2.6, and DS-V4-Pro tied at 87.5 Pretty impressive

Funny how Gemini is theoretically the best -- but in practice all the bugs in the interface mean I don't want to use it anymore. The worst is it forgets context (and lies about it), but it's very unreliable at reading pdfs (and lies about it). There's also no branch, so once the context is lost/polluted, you have to start projects over and build up the context from scratch again.

The sheer number of bugs and lack of meaningful improvements in Google products is a clear counterargument to the AI bull thesis

If AI was so good at coding, why can’t it actually make a usable Gemini/AI Studio app?

Re: DeepSeek v4

#254
post #43

Earlier quoted context omitted.

Where's the training data and training scripts since you are calling this open source ? Edit: it seems "open source" was edited out of the parent comment.

doesn't it get tiring after a while? using the same (perceived) gotcha, over and over again, for three years now? no one is ever going to release their training data because it contains every copyrighted work in existence. everyone, even the hecking-wholesome safety-first Anthropic, is using copyrighted data without permission to train their models. there you go.

Nvidia did with Nemo.

Re: DeepSeek v4

#255
post #197

Earlier quoted context omitted.

Weights are the source, training data is the compiler

Training data == source code, training algorithm == compiler, model weights == compiled binary.

Training algorithm is the programmer, weights are the code that you run in an interpreter

Re: DeepSeek v4

#256
post #143
post #63

I like the pelican I got out of deepseek-v4-flash more than the one I got from deepseek-v4-pro. https://simonwillison.net/2026/Apr/24/deepseek-v4/ Both generated using OpenRouter. For comparison, here's what I got from DeepSeek 3.2 back in December: https://simonwillison.net/2025/Dec/1/deepseek-v32/ And DeepSeek 3.1 in August: https://simonwillison.net/2025/Aug/22/deepseek-31/ And DeepSeek v3-0324 in March last year:…

DeepSeek pelicans are the angriest pelicans I’ve seen so far.

996 Pelican, lol

Re: DeepSeek v4

#257

> pricing "Pro" $3.48 / 1M output tokens vs $4.40 I’d like somebody to explain to me how the endless comments of "bleeding edge labs are subsidizing the inference at an insane rate" make sense in light of a humongous model like v4 pro being $4 per 1M. I’d bet even the subscriptions are profitable, much less the API prices. edit: $1.74/M input $3.48/M output on OpenRouter

And they actually say the prices will be "significantly" lower in second semester when Huawei 650 chips comes in.

Re: DeepSeek v4

#258
It is great! I asked the question what I always ask of new models ("what would Ian M Banks think about the current state of AI") and it gave me a brilliant answer! Funny enough the answer contained multiple criticisms of his own creators ("Chinese state entities", "Social Credit System").

Re: DeepSeek v4

#259

> pricing "Pro" $3.48 / 1M output tokens vs $4.40 I’d like somebody to explain to me how the endless comments of "bleeding edge labs are subsidizing the inference at an insane rate" make sense in light of a humongous model like v4 pro being $4 per 1M. I’d bet even the subscriptions are profitable, much less the API prices. edit: $1.74/M input $3.48/M output on OpenRouter

I mean, not one "bleeding edge" lab has stated they are profitable. They don't publish financials aside from revenue. And in Anthropic's case, they fuck with pricing every week. Clearly something is wrong here.

Re: DeepSeek v4

#260
post #96

MMLU-Pro: Gemini-3.1-Pro at 91.0 Opus-4.6 at 89.1 GPT-5.4, Kimi2.6, and DS-V4-Pro tied at 87.5 Pretty impressive

Funny how Gemini is theoretically the best -- but in practice all the bugs in the interface mean I don't want to use it anymore. The worst is it forgets context (and lies about it), but it's very unreliable at reading pdfs (and lies about it). There's also no branch, so once the context is lost/polluted, you have to start projects over and build up the context from scratch again.

I gave up on Gemini 3.1 Pro in VSCode after 2 hours. They fully refunded me.
Post reply on HN