Live data from Hacker News

DeepSeek R2 launch stalled as CEO balks at progress

reuters.com

81–90 of 186 posts

Re: DeepSeek R2 launch stalled as CEO balks at progress

#81

Earlier quoted context omitted.

I don't know why this isn't the crux of our current geopolitical spat. Surely it would be cheaper and easier for the CCP to develop their own chipmaking capacity than going to war in the Taiwan strait?

US will intervene militarily to stop China from taking control of TSMC if Taiwan isn't pressured by US to destroy the plants themselves, so I don't think taking Taiwan is a viable path to leading in silica only lowering US ability but given the current gap in GPUs it's not clear how helpful this is to China. So all in all I don't think China views taking Taiwan as beneficial in the AI race at all.

> US will intervene militarily

with a reality tv show dude being the commander in chief and a news reporter being the defense secretary.

life is tough in america, man.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#82
post #68

Earlier quoted context omitted.

As Russian I only recently started to understand that russian government was at wars for a lot of its existence from USSR times: https://en.wikipedia.org/wiki/List_of_wars_involving_Russia#... . Many invasions and wars in places Russia should have no business in. Most of them not publicized in the country. Unlike US it was not spreading liberal values of individual freedom and against violent dictatorships, actually…

> against violent dictatorships Then look up Latin America’s history, where the US actively worked to install and support such violent dictatorships. Some under the guise of protecting countries from the threat of communism - like Brazil, Argentina and Chile, and some explicitly to protect US company’s interests - like in Guatemala

> Chile

Yes fuckups happened. But then for results Russian intervention see CCP and how many people died from their hands and policies

Re: DeepSeek R2 launch stalled as CEO balks at progress

#83
post #26

Earlier quoted context omitted.

My experience with R1-0528 for python code generation was awful. But I was using a context length of 100k tokens, so that might be why. It scores decently in the lmarena code leaderboard, where context length is short.

Would love to see the system/user prompts involved, if possible. Personally I get it to write the same code I'd produce, which obviously I think is OK code, but seems other's experience differs a lot from my own so curious to understand why. I've iterated a lot on my system prompt so could be as easy as that.

The biggest reason I use Gemini is because it can still get stuff done at 100k context. The other models start wearing out at 30k and are done by 50k.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#84
post #80

Earlier quoted context omitted.

> Using H20 to serve DeepSeek V3 / R1 is just SUPER inefficient. Like, R1 is the most anti-H20 model released ever. Why? Any chance you have some links to read about why it’s the case?

MLA uses way more flops in order to conserve memory bandwidth, H20 has plenty of memory bandwidth and almost no flops. MLA makes sense on H100/H800, but on H20 GQA-based models are a way better option.

MLA as in multi-head latent attention?

Re: DeepSeek R2 launch stalled as CEO balks at progress

#85
post #50

Earlier quoted context omitted.

> if we were at a point where we'd have models developed and produced in the EU, as well as other parts of the world. But we have models developing and being produced outside of the US already, both in Asia but also Europe. Sure, it would be cool to see more from South America and Africa, but the playing field is not just in the US anymore, particularly when it comes to open weights (which seems more of a "world bene…

> when it comes to open weights (which seems more of a "world benefit" than closed APIs), then the US is lagging far behind. Llama (v4 notwithstanding) and Gemma (particularly v3) aren't my idea of lagging far behind...

> Llama (v4 notwithstanding) and Gemma (particularly v3) aren't my idea of lagging far behind...

While neat and of course Llama kicked off a large part of the ecosystem, so credit where credit is due, both of those suffer from "open-but-not-quite" as they have large documents of "Acceptable Use" which outlines what you can and cannot do with the weights, while the Chinese counter-parts slap a FOSS-compatible license on the weights and calls it a day.

We could argue if that's the best approach, or even legal considering the (probable) origin of their training data, but the end result remains the same, Chinese companies are doing FOSS releases and American companies are doing something more similar to BSL/hybrid-open releases.

It should tell you something when the legal department of one of these companies calls the model+weights "proprietary" while their marketing department continues to calling the same model+weights "open source". I know who I trust of those two to be more accurate.

I guess that's why I see American companies as being further behind, even though they do release something.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#86
post #26

Earlier quoted context omitted.

Would love to see the system/user prompts involved, if possible. Personally I get it to write the same code I'd produce, which obviously I think is OK code, but seems other's experience differs a lot from my own so curious to understand why. I've iterated a lot on my system prompt so could be as easy as that.

The biggest reason I use Gemini is because it can still get stuff done at 100k context. The other models start wearing out at 30k and are done by 50k.

The biggest reason I avoid Gemini (and all of Google's models I've tried) is because I cannot get them to produce the same code I'd produce myself, while with OpenAI's models it's fairly trivial.

There is something deeper in the model that seemingly can be steered/programmed with the system/user prompts and it still produces kind of shitty code for some reason. Or I just haven't found the right way of prompting Google's stuff, could also be the reason, but seemingly the same approach works for OpenAI, Anthropic and others, not sure what to make of it.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#87

Honestly, AI progress suffers because of these export restrictions. An open source model that can compete with Gemini Pro 2.5 and o3 is good for the world, and good for AI

> Honestly, AI progress suffers because of these export restrictions. An open source model that can compete with Gemini Pro 2.5 and o3 is good for the world, and good for AI

DeepSeek is not a charity, they are the largest hedge fund in China, nothing different from a typical wall street funds. They don't spend billions to give the world something open and free just because it is good.

When the model is capable of generating decent amount of revenues, or when there is conclusive evidence of showing being closed would lead to much higher profit, it will be closed.

Re: DeepSeek R2 launch stalled as CEO balks at progress

#88
post #9

"We had difficulties accessing OpenAI, our data provider." /s

Rumour was that DeepSeek used the outputs of the thinking steps in OpenAI's reasoning model (o1 at the time) to traing DeepSeek's Large Reasoning Model R1.

OpenAI used literally all available text owned by the entire human race to train o1/o3.

so what?

Re: DeepSeek R2 launch stalled as CEO balks at progress

#89
post #13

Earlier quoted context omitted.

Releasing the model has paid off handsomely with name recognition and making a significant geopolitical and cultural statement. But will they keep releasing the weights or do an OpenAI and come up with a reason they can't release them anymore? At the end of the day, even if they release the weights, they probably want to make money and leverage the brand by hosting the model API and the consumer mobile app.

If they continue to release the weights + detailed reports what they did, I seriously don't understand why. I mean it's cool. I just don't understand why. It's such a cut throat environment where every little bit of moat counts. I don't think they're naive. I think I'm naive.

If moving faster is a most, then open source AI could move faster than closed AI by not needing to be paranoid about privacy and welcoming external contributions

Re: DeepSeek R2 launch stalled as CEO balks at progress

#90
post #80

Earlier quoted context omitted.

MLA uses way more flops in order to conserve memory bandwidth, H20 has plenty of memory bandwidth and almost no flops. MLA makes sense on H100/H800, but on H20 GQA-based models are a way better option.

MLA as in multi-head latent attention?

Yes
Post reply on HN