Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

171–180 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#171

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

People and companies trust OpenAI and Anthropic, rightly or wrongly, with hosting the models and keeping their company data secure. Don't underestimate the value of a scapegoat to point a finger at when things go wrong.

But they also trust cloud platforms like GCP to host models and store company data.

Why would a company use an expensive proprietary model on Vertex AI, for example, when they could use an open-source one on Vertex AI that is just as reliable for a fraction of the cost?

I think you are getting at the idea of branding, but branding is different from security or reliability.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#172

Earlier quoted context omitted.

Oh they need control of models to be able to censor and ensure whatever happens inside the country with AI stays under their control. But the open-source part? Idk I think they do it to mess with the US investment and for the typical open source reasons of companies: community, marketing, etc. But tbh especially the messing with the US, as a european with no serious competitor, I can get behind.

This is the rare earth minerals dumping all over again. Devalue to such a price as to make the market participants quit, so they can later have a strategic stranglehold on the supply. This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it. It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may beco…

Are you by chance an OpenAI investor?

We should all be happy about the price of AI coming down.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#173
post #43
post #24

Earlier quoted context omitted.

1. Have you seen the Qwen offerings? They have great multi-modality, some even SOTA.

Qwen Image and Image Edit were among the best image models until Nano Banana Pro came along. I have tried some open image models and can confirm , the Chinese models are easily the best or very close to the best, but right now the Google model is even better... we'll see if the Chinese catch up again.

I'd say Google still hasn't caught up on the smaller model side at all, but we've all been (rightfully) wowed enough by Pro to ignore that for now.

Nano Banano Pro starts at 15 cents per image at Add in the power of fine-tuning on their open weight models and I don't know if China actually needs to catch up.

I finetuned Qwen Image on 200 generations from Seedream 4.0 that were cleaned up with Nano Banana Pro, and got results that were as good and more reliable than either model could achieve otherwise.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#174

Earlier quoted context omitted.

Oh they need control of models to be able to censor and ensure whatever happens inside the country with AI stays under their control. But the open-source part? Idk I think they do it to mess with the US investment and for the typical open source reasons of companies: community, marketing, etc. But tbh especially the messing with the US, as a european with no serious competitor, I can get behind.

This is the rare earth minerals dumping all over again. Devalue to such a price as to make the market participants quit, so they can later have a strategic stranglehold on the supply. This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it. It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may beco…

> cheat on the fair market

Can you really view this as a cheat this when the US is throwing a trillion dollars in support of a supposedly "fair market"?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#175

Earlier quoted context omitted.

Either... Better (UX / ease of use) Lock in (walled garden type thing) Trust (If an AI is gonna have the level of insight into your personal data and control over your life, a lot of people will prefer to use a household name)

> Trust (If an AI is gonna have the level of insight into your personal data and control over your life, a lot of people will prefer to use a household name. Not Google, and not Amazon. Microsoft is a maybe.

People trust google with their data in search, gmail, docs, and android. That is quite a lot of personal info, and trust, already.

All they have to do is completely switch the google homepage to gemini one day.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#176
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

>winning on cost-effectiveness Nobody is winning in this area until these things run in full on single graphics cards. Which is sufficient compute to run even most of the complex tasks.

Why does that matter? They wont be making at home graphics cards anymore. Why would you do that when you can be pre-sold $40k servers for years into the future

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#177
post #153

Earlier quoted context omitted.

on what hypothetical grounds would you be more meaningfully able to sue the american maker of a self-hosted statistical language model that you select your own runtime sampling parameters for after random subtle security vulnerabilities came out the other side when you asked it for very secure code? put another way, how do you propose to tell this subtle nefarious chinese sabotage you baselessly imply to be commonpla…

[flagged]

sorry, is your contention here "spurious accusations don't require evidence when aimed at designated state enemies"? because it feels uncharitably rude to infer that's what you meant to say here, but i struggle to parse this in a different way where you say something more reasonable.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#178

Earlier quoted context omitted.

"Mixture-of-experts", AKA "running several small models and activating only a few at a time". Thanks for introducing me to that concept. Fascinating. (commentary: things are really moving too fast for the layperson to keep up)

that's not really a good summary of what MoEs are. you can more consider it like sublayers that get routed through (like how the brain only lights up certain pathways) rather than actual separate models.

The gains from MoE is that you can have a large model that's efficient, it lets you decouple #params and computation cost. I don't see how anthropomorphizing MoE brain affords insight deeper than 'less activity means less energy used'. These are totally different systems, IMO this shallow comparison muddies the water and does a disservice to each field of study. There's been loads of research showing there's redundancy in MoE models, ie cerebras has a paper[1] where they selectively prune half the experts with minimal loss across domains -- I'm not sure you could disable half the brain and notice a stupefying difference.

[1] https://www.cerebras.ai/blog/reap

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#180

Earlier quoted context omitted.

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…

If the Chinese model becomes better than competitors, these worries will suddenly disappear. Also, there are plenty startups and enterprises that are running fine-tuned versions of different OS models.

No… Nobody I work for will touch these models. The fear is real that they have been poisoned or have some underlying bomb. Plus y’know, they’re produced by China, so they would never make it past a review board in most mega enterprises IME.
Post reply on HN