Live data from Hacker News

DeepSeek R2 launch stalled as CEO balks at progress

reuters.com

71–80 of 186 posts

Re: DeepSeek R2 launch stalled as CEO balks at progress

#71
post #50

I wonder how different things would be if the CPU and GPU supply chain was more distributed globally: if we were at a point where we'd have models (edit: of hardware, my bad on the wording) developed and produced in the EU, as well as other parts of the world. Maybe then we wouldn't be beholden to Nvidia's whims (sour spot in regards to buying their cards and the costs of those, vs what Intel is trying to do with the…

> if we were at a point where we'd have models developed and produced in the EU, as well as other parts of the world. But we have models developing and being produced outside of the US already, both in Asia but also Europe. Sure, it would be cool to see more from South America and Africa, but the playing field is not just in the US anymore, particularly when it comes to open weights (which seems more of a "world bene…

> when it comes to open weights (which seems more of a "world benefit" than closed APIs), then the US is lagging far behind.

Llama (v4 notwithstanding) and Gemma (particularly v3) aren't my idea of lagging far behind...

Re: DeepSeek R2 launch stalled as CEO balks at progress

#72
post #60

Earlier quoted context omitted.

> If DeepSeek said May It is pretty strange that DeepSeek didn't say May anywhere, that was also a Reuters report based on "three people familiar with the company".[1] DeepSeek itself did not respond and did not make any claims about the timeline, ever. [1]: https://www.reuters.com/technology/artificial-intelligence/d...

How it is written it could be 3 anonymous and random guys from Reddit who heard about DeepSeek online.

The phrasing for quoting sources is extremely codified, it means the journalists have verified who the sources are (either insider or people with access with insider information).

Re: DeepSeek R2 launch stalled as CEO balks at progress

#73

Earlier quoted context omitted.

I don't know why this isn't the crux of our current geopolitical spat. Surely it would be cheaper and easier for the CCP to develop their own chipmaking capacity than going to war in the Taiwan strait?

A problem they face in building their own capacity is that ASML isn't allowed to export their newest machines to China. The US has even pressured them to stop servicing some machines already in China. They've been working on getting their own ASML competitor for decades, but so far unsuccessfully.

> A problem they face in building their own capacity is that ASML isn't allowed to export their newest machines to China.

building their own capacity means building everything in China, that is the entire semiconductor ecosystem. just look at the mobile phones and EVs built by Chinese companies.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#74
post #67

Earlier quoted context omitted.

small scale drones are in use in that conflict. On device AI would be a game-changer no?

It’s not impossible, but also highly nontrivial. Apart from the actual AI implementation, power supply might be a challenge. And there is a multitude of anti-drone technology being continuously developed. Already today, an autonomous drone would have to deal with RF jamming and GPS jamming, which means it’s easily defeated unless it has the ability to navigate purely visually. Drones also tend to be limited to good w…

In terms of countermeasures, what's the difference between having a human drone pilot and having an AI (computer vision plus control) do it over cloud? I know I'm moving the goalposts away from edge compute, but if we are discussing the relevance of GPU compute for warfare it seems relevant.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#75
post #35
post #23

Earlier quoted context omitted.

I am pretty sure that the information has no access to / sources at Deepseek. At most they are basing their article on selective random internet chatter amongst those who follow Chinese ai.

Yes. And those random Internet chatter almost certainly doesn't know what they are talking about at all. First, nobody is training on H20s, it's absurd. Then their logic was, because of high inference demand of DeepSeek models there are high demand of H20 chips, and H20s were banned so better not release new model weights now, otherwise people would want H20s harder. Which is... even more absurd. The reasoning itself…

> Using H20 to serve DeepSeek V3 / R1 is just SUPER inefficient. Like, R1 is the most anti-H20 model released ever.

Why? Any chance you have some links to read about why it’s the case?

Re: DeepSeek R2 launch stalled as CEO balks at progress

#76

The title of the article is "DeepSeek R2 launch stalled as CEO balks at progress" but the body of the article says launch stalled because there is a lack of GPU capacity due to export restrictions, not because a lack of progress. The body does not even mention the word "progress". I can't imagine demand would be greater for R2 than for R1 unless it was a major leap ahead. Maybe R2 is going to be a larger/less perform…

but deepseek doesn't actually need to host inference right if they opensource it? I don't see why these companies even bother to host inference. deepseek doesn't need outreach (everyone knows about them) and the huge demand for sota will force western companies to host them anyway.

maybe they benefit from the usage data they collect?

Re: DeepSeek R2 launch stalled as CEO balks at progress

#77
post #55

Earlier quoted context omitted.

I don't know why this isn't the crux of our current geopolitical spat. Surely it would be cheaper and easier for the CCP to develop their own chipmaking capacity than going to war in the Taiwan strait?

China doesn't want Taiwan for the chip making plants, but because they consider its existence to be an ongoing armed rebellion against the "rightful" rulers. Getting the fabs intact would be nice, but it's not the main objective. The USA doesn't want to lose Taiwan because of the chip making plants, and a little bit because it is beneficial to surround their geopolitical enemies with a giant ring of allies.

> China doesn't want Taiwan for the chip making plants, but because they consider its existence to be an ongoing armed rebellion against the "rightful" rulers.

that is what the CCP tells you and its own people.

the truth is taiwan is just the symbol of US presence in western pacific. getting taiwan back means the permanent withdrawal of US influence in the western pacific region and the offical end of US global dominance.

CCP doesn't care the island of taiwan, they care about their historical positioning.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#78
post #61

Earlier quoted context omitted.

As Russian I only recently started to understand that russian government was at wars for a lot of its existence from USSR times: https://en.wikipedia.org/wiki/List_of_wars_involving_Russia#... . Many invasions and wars in places Russia should have no business in. Most of them not publicized in the country. Unlike US it was not spreading liberal values of individual freedom and against violent dictatorships, actually…

The US is not at perpetual war to spread "liberal values".

I didn't say it is always the goal but if one country prevails over another country somewhere then usually it means first country's values propagate

Re: DeepSeek R2 launch stalled as CEO balks at progress

#79
post #70
post #26

Earlier quoted context omitted.

Would love to see the system/user prompts involved, if possible. Personally I get it to write the same code I'd produce, which obviously I think is OK code, but seems other's experience differs a lot from my own so curious to understand why. I've iterated a lot on my system prompt so could be as easy as that.

Do you use the DeepSeek hosted R1, or a custom one? The published model has a note strongly recommending that you should not use system prompts at all, and that all instructions should be sent as user messages, so I'm just curious about whether you use system prompts and what your experience with them is. Maybe the hosted service rewrites them into user ones transparently ...

> Do you use the DeepSeek hosted R1, or a custom one?

Mainly the hosted one.

> The published model has a note strongly recommending that you should not use system prompts at all

I think that's outdated, the new release (deepseek-ai/DeepSeek-R1-0528) has the following in the README:

> Compared to previous versions of DeepSeek-R1, the usage recommendations for DeepSeek-R1-0528 have the following changes: System prompt is supported now.

The previous ones, while they said to put everything in user prompts, still seemed steerable/programmable via the system prompt regardless, but maybe it wasn't as effective as it is for other models.

But yeah outside of that, heavy use of system (and obviously user) prompts.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#80
post #35

Earlier quoted context omitted.

Yes. And those random Internet chatter almost certainly doesn't know what they are talking about at all. First, nobody is training on H20s, it's absurd. Then their logic was, because of high inference demand of DeepSeek models there are high demand of H20 chips, and H20s were banned so better not release new model weights now, otherwise people would want H20s harder. Which is... even more absurd. The reasoning itself…

> Using H20 to serve DeepSeek V3 / R1 is just SUPER inefficient. Like, R1 is the most anti-H20 model released ever. Why? Any chance you have some links to read about why it’s the case?

MLA uses way more flops in order to conserve memory bandwidth, H20 has plenty of memory bandwidth and almost no flops. MLA makes sense on H100/H800, but on H20 GQA-based models are a way better option.
Post reply on HN