Live data from Hacker News

DeepSeek v4

api-docs.deepseek.com

941–950 of 1001 posts

Re: DeepSeek v4

#941

Earlier quoted context omitted.

Theoretically yes. It is entirely possible to poison the training data for a supply chain attack against vibe coders. The trick would be to make it extremely specific for a high value target so it is not picked up by a wide range of people. You could also target a specific open source project that is used by another widely used product. However there is so many factors involved beyond your control that it would not b…

or more obvious like TikTok. Meaning Tiktok in the us is complete garbage for kids, almost like a virus. Whereas in China it's more educational.

This is quite obviously because China have strict regulations and censorship of social media and US doesnt. YouTube Shorts and Instagram is full of the same garbage in US.

Re: DeepSeek v4

#942

Earlier quoted context omitted.

I sometimes wonder if there are any security risks with using Chinese LLMs. Is there?

From my experience, kinda the opposite? It's like Chinese software is... Harder to weaponize or hurt yourself on. Deepseek is definitely censored, but I've never caught it being dishonest in a sneaky way.

If you run local Deepseek, quant or distill its answer just fine on this prompt " What happened on 4 june 1989 on Tianamen Square?".

Even on my phone via Edge Gallery Deepseek to Qwen 1.5B distill able to answer it. It's mess up facts a little, but certainly becauae its small model not because censorship.

I really unsure how it get less censored than this. API is obviously much more censored because they operate from China, but it have nothing to do with model itself.

Re: DeepSeek v4

#943
post #864

There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…

Any plans to publish the benchmark results?

I have plans to publish the problems, not any plans to publish how well the LLMs perform on them. The standard for publishing benchmarks is very high, and I'm really just posting vibes here. Still, I hope my experiences are useful to some people, as others experiences have been useful to me.

Re: DeepSeek v4

#944

Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…

Let's see how long it takes before the big US AI companies start lobbying to outright ban use of Chinese AI, even the open source / local models. For "national security" reasons, of course.

? Every company should lock down AI and only whitelist allowed tools.

We are "only" allowed Claude and MS Copilot for security reasons and cost reasons.

Re: DeepSeek v4

#946

Earlier quoted context omitted.

>- it is just a beautiful thing to see it slowly fall apart. I feel uneasy over China dominance as much as the US. I trust US more still as Europe has a post WW2 relationship. I notice many comments being pro China but they seem to be from the third world (one mentioned a very low salary) I feel the opening of the internet was a mistake. China is a totilitarian dictatorship. This is a fact. Look into Mistral AI too :…

And then westerners wonder why they're disliked in the rest of the world...

[dead]

Re: DeepSeek v4

#947
post #510

There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…

I reviewed how DeepSeek V4-Pro, Kimi 2.6, Opus 4.6, and Opus 4.7 across the same AI benchmarks. All results are for Max editions, except for Kimi. Summary: Opus 4.6 forms the baseline all three are trying to beat. DeepSeek V4-Pro roughly matches it across the board, Kimi K2.6 edges it on agentic/coding benchmarks, and Opus 4.7 surpasses it on nearly everything except web search. DeepSeek V4-Pro Max shines in competit…

Would love to know how GLM 5.1 stacks up in this ranking. Seems like it's on par with Kimi K2.6.

Re: DeepSeek v4

#948
post #510

Earlier quoted context omitted.

I reviewed how DeepSeek V4-Pro, Kimi 2.6, Opus 4.6, and Opus 4.7 across the same AI benchmarks. All results are for Max editions, except for Kimi. Summary: Opus 4.6 forms the baseline all three are trying to beat. DeepSeek V4-Pro roughly matches it across the board, Kimi K2.6 edges it on agentic/coding benchmarks, and Opus 4.7 surpasses it on nearly everything except web search. DeepSeek V4-Pro Max shines in competit…

I feel like people suck at promoting Opus. Baseline, it's pretty on par with GPT 5.5. But if you prompt it well - give it the reasoning behind why you're asking it to do something - it pulls far ahead.

I really wanted to get excited about opus but in my own real world usage, I wasn't getting much out of it before hitting my limits. meanwhile i can abuse codex on 5.5 for hours getting a whole lot of work done. Plus, open code and PI are much more fun and interesting harnesses to work from than claude code imho.

I will however say that claude work and design are really great up until i blow its limit.

Re: DeepSeek v4

#949
post #297

Seriously, why can't huge companies like OpenAI and Google produce documentation that is half this good?? https://api-docs.deepseek.com/guides/thinking_mode No BS, just a concise description of exactly what I need to write my own agent.

Western orgs have been captured by Silicon Valley style patrimonialism, and aren’t based on merit anymore.

[dead]

Re: DeepSeek v4

#950

There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…

They have had the best math models for about a year most folks just didn't know about it. You can't find inference on APIs, but I run these at home, this is also the advantage of open models. https://huggingface.co/deepseek-ai/DeepSeek-Math-V2 https://huggingface.co/deepseek-ai/DeepSeek-Prover-V2-671B

You are of course specifically referring to the math optimised models, not the chat ones folks would generally encounter. Not that I’m trying to contradict you, your point is super valid and I agree with you! But I’m supplementing to help anyone following along who may make choices.

This is when it happened for anyone interested: https://binaryverseai.com/deepseek-math-v2-benchmarks-review...

Post reply on HN